Multimodal ai

Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs (Ep. 1018)

Alibaba's Qwen3.8-Max: Open-Weight Model Surpasses Most American Frontier Labs (Ep. 1018)

Jon Krohn dissects Alibaba's Qwen 3.8 Max, a 2.4-trillion-parameter Mixture-of-Experts (MoE) model positioned as the largest open-weight release in history if its promised weights ship. The discussion covers its multimodal capabilities, 1M token context window, and performance competitive with Anthropic's Claude Fable 5. Key highlights include its advanced multi-day agentic capabilities and aggressively low pricing ($2 in / $6 out per million tokens), intensifying the AI price war. Krohn also provides critical insights into the safety of using Chinese models, emphasizing data handling practices and the benefits/risks across different deployment scenarios.

The New Primitives: Building AI Native Software — Kwindla Kramer, Daily

The New Primitives: Building AI Native Software — Kwindla Kramer, Daily

Kwindla Hultman Kramer argues that current AI agents are akin to 1995 web pages – a foundational primitive, not the ultimate destination. Drawing a historical parallel, he predicts the emergence of "AI native software" that will build upon and surpass agents, much like web applications evolved from simple web pages. He illustrates this through computing history, highlighting transformative shifts like VisiCalc's impact on accounting and the vision of Apple's Knowledge Navigator, and concludes by showcasing a game, Gradient Bang, that demonstrates the core primitives of this future AI native software.

Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song

Agents at Scale: Inside MiniMax's Model and the Infrastructure Behind It — Olive Song

Olive Song, RL lead at MiniMax, details the engineering behind MiniMax's open-weight models, focusing on M3's multimodal and agentic capabilities, the necessity of day-zero inference stack readiness, and continuous GPU kernel optimization. She discusses multimodal training challenges, long-horizon task evaluation, and expresses optimism for open models rapidly closing the gap with frontier labs.

Your AI Evals Are Lying

Your AI Evals Are Lying

Andrew Burt of Luminos discusses how current AI risk evaluation methods are insufficient, advocating for a "high dimensionality" approach using granular sub-risks and diverse legal and technical expertise. He critiques common practices like guardrails and single-LLM evaluations, highlighting the need for multimodal systems and continuous, automated monitoring to address the evolving complexities of AI, particularly with the rise of open-weight models.

Why Physical AI Is the Next Platform Shift

Why Physical AI Is the Next Platform Shift

Encord Co-CEO Eric Landau reflects on his transition from a lucrative quant career to founding an AI startup, driven by a deep belief in AI's paradigm-shifting potential. He discusses Encord's slow, compounding path to product-market fit, the pivotal role of Physical AI, and the importance of embracing the emotional rollercoaster of startup life.

Perception Agents — Antje Barth, Amazon AGI Lab

Perception Agents — Antje Barth, Amazon AGI Lab

Antje Barth of Amazon AGI Lab discusses the architectural gap in current AI agents, which excel at individual tasks but fail at complex, end-to-end workflows due to a lack of reliability and contextual understanding. She introduces "Perception Agents"—AI systems that see, reason, and act on computers like humans, using visual and multimodal input to enable reliable collaboration and close the perception-action loop, highlighting new open-source tools for annotation and verification.