Foundation models

Multimodal & Embodied Intelligence (Pt 1), Panel on Multimodal AI: Progress, Pitfalls, Possibilities

Multimodal & Embodied Intelligence (Pt 1), Panel on Multimodal AI: Progress, Pitfalls, Possibilities

This session explored Multimodal and Embodied Intelligence, featuring talks on hybrid AI in robotics (classical vs. end-to-end), AI's role in healthcare (focusing on NCDs, deployment, and uncertainty modeling), and fundamental perception challenges in multimodal reasoning (using educational video QA and visual puzzles). A panel discussed the impact of foundation models, the blurred lines between AGI and human-like AI, critical deployment pitfalls (human factors, efficiency, architectural limits), and future directions, emphasizing task-specific models and the redefinition of 'foundation models.'

Plenary Talk 3: Challenges and Research Opportunities for Global Hyperscale Services

Plenary Talk 3: Challenges and Research Opportunities for Global Hyperscale Services

This talk provides a comprehensive overview of cellular aging, brain function, and the mechanisms of cognitive decline, particularly focusing on Alzheimer's disease. It delves into the role of various cell types, neural communication, and methods for assessing cognition. The speaker highlights research challenges, the limitations of current pharmacological interventions, and the critical importance of non-pharmacological lifestyle interventions. A significant portion of the discussion is dedicated to the Centre for Brain Research's (CBR) multi-disciplinary efforts in India, including large-scale cohort studies, multimodal data collection, and the development of AI-driven tools and a localized foundation model for the Indian brain to address neurodegeneration.

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

How Cursor Trained Composer on Fireworks: Distributed Infrastructure for High-Performance RL

Cursor's Federico Cassano and Fireworks' Dmytro Dzhulgakov detail their collaboration on Composer 2, a specialized foundation model for software engineering. They discuss their top-down training strategy, the infrastructure challenges of large-scale distributed Reinforcement Learning on sparse models, and how model specialization achieves frontier performance with superior efficiency.

End-to-End Foundation Models for the Energy Industry — with Jazmia Henry

End-to-End Foundation Models for the Energy Industry — with Jazmia Henry

Jazmia Henry details the end-to-end process of building specialized foundation models for the energy industry. She covers the four key stages from data curation of unstructured, handwritten documents to optimizing inference, and introduces her Grounded Continuous Evaluation (GCE) framework to combat reward hacking in reinforcement learning.

How Transformers Finally Ate Vision – Isaac Robinson, Roboflow

How Transformers Finally Ate Vision – Isaac Robinson, Roboflow

Isaac Robinson from Roboflow explains why Vision Transformers (ViTs), despite their initial disadvantages in computational complexity and lack of inductive bias, ultimately surpassed Convolutional Neural Networks (CNNs) for computer vision tasks. The talk covers the critical roles of massive, ViT-specific pre-training methods like MAE and DINO, the architectural evolution through models like Swin, ConvNeXt, and Hiera, and optimizations borrowed from the LLM ecosystem. It culminates in a discussion on the practical deployment challenges of large foundation models like SAM and how Neural Architecture Search can bridge the gap.

Waymo's Dmitri Dolgov: 20 Million Rides and the Road to Full Autonomy

Waymo's Dmitri Dolgov: 20 Million Rides and the Road to Full Autonomy

Dmitri Dolgov, co-CEO of Waymo, discusses the 20-year journey from the DARPA challenge to full autonomy. He explains the Waymo Foundation Model—a multimodal world action model powering the driver, simulator, and critic—and how their "end-to-end plus" architecture enables superhuman safety and exponential scaling.