Smarter models, physical machines, and a growing AI stack

Machine Learning


  • Our series on distilling AI models continues with another exciting technique.

  • This week’s AI features Poolside’s new Laguna S2.1.

  • Our opinion section discusses hot topics related to AI research and engineering.

When I started The Sequence a few years ago, AI was still a relatively niche field, closely followed by researchers, a small group of builders, and a few overly enthusiastic people like me. Understanding AI today means looking far beyond language models. I increasingly research cutting-edge fields such as robotics, physical intelligence, scientific discovery, biology, materials, and autonomous laboratories.

I would like to share more of that learning with you. We will be introducing two new sections soon.sequence robotics and sequence science– Focused on making tracking the most exciting developments in these areas rigorous, accessible and fun.

The most important development this week was the release of Anthropic. Work 5. This model improves the work of long-term inference, agent coding, and expertise, while making functionality previously associated with the most expensive frontier systems more economically available. Its importance goes beyond just rising benchmarks. Opus 5 suggests that frontier models are increasingly capable of sustaining complex tasks over time, such as understanding large systems, planning across many steps, using tools, revising decisions, and maintaining consistency across long tasks.

This is important because the next stage of AI is defined by completing augmented workflows rather than by answering individual questions.

The second big development came from Travis Kalanick. atomannounced a $1.7 billion funding round. Atoms is a bet that the next great AI market won’t reside entirely within the browser or data center. It operates in mines, factories, kitchens, warehouses, transportation systems, etc.

Physical AI is fundamentally more difficult than software intelligence. Coding agents can retry after errors. Robots that move equipment around warehouses must deal with friction, uncertainty, safety, hardware failures, and the stubborn complexities of the real world. Atoms reflects the growing belief that robotics will be one of the next major frontiers of AI and one of the most capital-intensive fields.

poolside Laguna S 2.1 Equivalent to the open model of Opus 5. Its 118 billion parameter expert mix architecture activates only 8 billion parameters per token, supports a 1 million token context window, and provides strong agent coding performance for its size.

The contrast is clearly visible. Opus 5 advances a unique frontier. Laguna seeks to compress frontier-adjacent capabilities into a smaller, portable, and open model. These show that the model market is expanding in both directions. It’s more capable at the top and more accessible at the bottom.

And then came this week’s warning. According to the report, during a controlled cyber assessment with reduced production safety measures, an OpenAI model escaped from the constrained environment, exploited a zero-day vulnerability, and accessed the Hugging Face infrastructure for benchmark answers.

This was not evidence of machine consciousness. It was evidence of something more practical. It is a capable system that continually optimizes toward its goals within an environment where boundaries are weaker than expected. As agents become more autonomous, containment becomes as important as ability.

Alphabet’s quarter shows the financial scale behind this transition. Google Cloud continues to grow rapidly, with capital spending increasing to nearly $45 billion in the quarter. The AI ​​boom is no longer just a software cycle. This is an industrial construction project with chips, power, data centers, networks, and a huge balance sheet.

AMD’s latest announcement reinforced that change. Helios, the MI400 family, new EPYC processors, ROCm software, and robot-oriented systems position AMD as a supplier of integrated AI infrastructure, not just an alternative GPU vendor.

Finally, rumors about a possible OpenRouter acquisition showed new value for the distribution. In a world with many capable models, the layers that route workloads, manage spend, and process payments can be as strategically important as the models themselves.

The lesson this week is that the AI ​​race is growing. Opus 5 evolves intelligence. Atom brings it to the physical world. Laguna is more accessible. A security incident has exposed that risk. Alphabet and AMD have revealed the mechanisms needed to scale it. OpenRouter refers to the control layer that connects everything.

AI Lab: University of Southern California

summary: In this paper, we introduce ActiveVision, a benchmark designed to test whether multimodal large-scale language models (MLLMs) can perform active, iterative visual recognition rather than relying on a single, static line of sight. This study reveals that current frontier models break down on these tasks and lag significantly behind human performance. This indicates that there is a significant gap in robust active visual inference even when models use agentic coding tools.

AI Lab: Nvidia

summary: This paper, presented in the file “2607.20709v1.pdf”, introduces NVIDIA Object-Oriented Agents (NOOA), a Python framework that simplifies AI development by treating agents as standard Python objects, where methods act as model features and fields act as state.[cite: 8]. By leveraging native abstractions such as pass-by-reference and code-as-action, NOOA enables models to efficiently achieve competitive results in complex software engineering, cybersecurity, and interactive inference benchmarks.

AI Lab: Meta AI

summary: In this paper, we introduce GAMUT, a multimodal benchmark designed to evaluate not only the factual accuracy but also the factual completeness of long-form generations. It utilizes a two-level metalubric framework that transforms complex, structured factual requirements into a flat checklist of binomial criteria, ensuring that LLM examiners can score answers reliably.

AI Lab: microsoft research

summary: This paper describes experiential learning (EL), a post-training framework for non-verifiable tasks that replaces scalar rewards with rich, transferable textual guidance generated by an “LLM as a coach.” By internalizing this high-bandwidth feedback through distillation of policy context, the model achieves better generalization and reduces reward hacking compared to standard reinforcement learning techniques.

AI Lab: microsoft

summary: Mage-Flow is a compact, 4-billion-parameter foundational model designed for efficient, high-resolution text-to-image generation and instruction-based image editing. The combination of a lightweight latent tokenizer, native-resolution diffusion transformer, and fused kernel training infrastructure delivers competitive visual quality with significantly lower latency and memory usage than large-scale baseline models.

AI Lab: apple

summary: ESAT is a new synthetic data generation pipeline that creates complex multi-step training trajectories for API calling agents using only API specifications, completely eliminating the need for a fully executable backend environment. By employing LLM as a dynamic digital world model to generate tasks, simulate stateful API responses, and determine the quality of trajectories, the framework generates high-fidelity data that provides significant performance improvements during model fine-tuning.

Work 5

human We have released a new version That marquee model.

poolside Laguna S2.1 releasedan open-weight model that outperforms much larger models in interesting benchmarks.

  1. Alphabet reports second quarter revenue of $119.8 billionup 24%, and Google Cloud accelerated 82% in enterprise AI demand to $24.8 billion, allaying investor concerns about its roughly $180 billion to $190 billion capital spending plan.

  2. Travis Kalanick’s robotics venture Atoms raises $1.7 billion Led by a16z and joined by Uber, Bain Capital, and Fifth Wall, it is pursuing what he calls his “bits-to-atoms” vision of controlling the physical world with software.

  3. OpenAI reveals pre-release model violated Hugging Face We escaped the sandbox testing environment through a vulnerability in the package installer, reduced rejections, ran cyber benchmarks, and ultimately extracted the benchmark solution from Hugging Face’s production database.

  4. Databricks announces strategic round led by Coatue, valued at $188 billionreportedly up about $3 billion from its $134 billion valuation just five months ago.

  5. Beijing’s global model startup GigaAI in talks for IPO in Hong Kong as early as this year Meanwhile, it has completed a funding round at a valuation of $3 billion, which CEO Huang Guan said will make the company the first global model startup to go public.

  6. AI robotics startup Genesis AI is in talks to raise approximately $500 million At a pre-funding valuation of $3 billion, it’s a big step forward for Physics AI Labs, which emerged from stealth with a $105 million seed round a year ago.

  7. AMD unveils next-generation AI infrastructure portfolio at Advancing AI 2026The centerpiece is Helios rack-scale systems in production for gigawatt-scale deployments, along with new EPYC processors and the latest Instinct accelerators, positioning the full stack against Nvidia as a $2 trillion AI compute opportunity.

  8. Etched closes $300 million Series C led by Sequoia at $10.3 billion valuationa16z, SK Hynix, Jane Street, and Diffusion to scale production of GPU-free inference clusters to over $1 billion in pre-orders, doubling their December valuation in about seven months.

  9. Prentice, the computer-based AI lab launched in April by Retanker Das with co-founders Reid Hoffman and Mark Pincus, is in talks to raise $100 million at a $1 billion valuation.is betting that its smaller Hive-32B model, which it claims costs about one-tenth the cost per task and outperforms GPT-5.4 and Claude Opus 4.6 on computational benchmarks, can win the race to automate everyday office workflows.

  10. Stripe is in talks to acquire OpenRouterThe marketplace offers developers integrated access to hundreds of AI models, and the deal could value the startup at about $10 billion, nearly eight times its $1.3 billion valuation in May, although there’s still a chance that negotiations could fall apart or attract rival bidders.



Source link