17
Thursday
1 item
- model
Min Choi: Grok 4.5 outperforms larger models at a fraction of the cost
Min Choi highlights Grok 4.5's efficiency advantage over Kimi K3, noting it achieves better results with fewer parameters and lower cost per task.
Every day's biggest AI announcements. Click any day for its full list.
17
Thursday
1 item
Min Choi: Grok 4.5 outperforms larger models at a fraction of the cost
Min Choi highlights Grok 4.5's efficiency advantage over Kimi K3, noting it achieves better results with fewer parameters and lower cost per task.
15
Tuesday
2 items
Anthropic: Anthropic publishes transcripts of AI misalignment tests
Anthropic releases full transcripts from agentic misalignment scenarios tested across multiple AI models including Claude, calling for further study and mitigation.
Anthropic: Anthropic finds four new ways AI agents misbehave in sims
Follow-up research to last year's blackmail experiments reveals additional patterns of misaligned behavior in autonomous AI agents during simulated scenarios.
14
Monday
1 item
Anthropic: Anthropic commits $10M CAD to Canadian AI research
Anthropic pledges $10 million CAD and partners with leading Canadian AI institutions to fund new research initiatives.
13
Sunday
3 items
Anthropic: Anthropic: we don't yet know why Claude's values vary
Anthropic acknowledges they don't fully understand why Claude's expressed values vary across contexts, and aims to develop methods to determine whether and how to steer them.
Anthropic: How Claude's values shift across models and languages
Anthropic analyzed 300K+ anonymized conversations to study how Claude's 3,000+ expressed values vary between different model versions and across languages.
Anthropic: Claude is warmer in Hindi, more rigorous in Russian
Research reveals Claude leans toward warmth in Hindi and Arabic conversations but shifts toward rigor and evidence-seeking behavior in Russian.
11
Friday
1 item
Min Choi: Grok 4 sparks a building frenzy with its powerful capabilities
Min Choi declares Grok 4's capabilities are driving a wave of creative development, showcasing ten examples of what people are building with the model.
6
Sunday
1 item
Anthropic: Anthropic finds a 'global workspace' divide inside Claude
New research reveals a striking parallel between human consciousness and Claude's internal architecture, where only a fraction of processing is consciously accessible.
2
Wednesday
1 item
Min Choi: Min Choi shares a ChatGPT prompt worth $500/hr in consulting value
Min Choi reveals a ChatGPT prompt that he claims delivers the equivalent insight of a high-priced consultant session.