New human AI models target coding and enterprise operations
Anthropic announced Claude Opus 4.6, which introduces a 1 million-token context window and automatic agent tuning capabilities, as the AI company looks to expand beyond software development into broader enterprise applications.
The San Francisco-based company said the model offers improved performance for coding tasks, financial analysis, and document processing compared to previous models. Anthropic positions this release as strengthening its position in enterprise AI workflows, an increasingly crowded market in which it competes directly with OpenAI and Google.
“We are committed to building the most capable, reliable and secure AI systems,” an Anthropic spokesperson said. “Opus 4.6 offers even better planning and helps you solve your most complex coding tasks.”
The release comes three days after OpenAI announced its desktop application for the Codex AI coding system, highlighting the rapid pace of competition in AI development tools. Anthropic announced in November that its coding product, Claude Code, reached $1 billion in annual revenue six months after becoming generally available.
Extended context and agent coordination
Opus 4.6 supports up to 1 million contextual tokens in the beta version of Anthropic’s developer platform. This is a significant increase from the 200,000 token limit in previous Opus versions. This enhancement allows the model to handle larger codebases and longer documents without having to split tasks into multiple requests.
The company also introduced Agent Teams in Claude Code as a research preview, allowing multiple AI agents to work on segmented parts of a project simultaneously. Scott White, head of product at Anthropic, likened this feature to coordinating a team of humans working in parallel.
Anthropic says Opus 4.6 addresses context degradation, a common issue where AI performance degrades as conversations get longer. On a search benchmark that hides information in large amounts of text, the Sonnet 4.5 model scored 18.5%, while the Opus 4.6 scored 76%.
This model supports output of up to 128,000 tokens. Anthropic introduced adaptive thinking, which allows models to decide when to apply deeper inference, and four effort settings that developers can adjust to balance performance, speed, and cost.
benchmark performance
Anthropic reported that Opus 4.6 led in Terminal-Bench 2.0, an evaluation of AI agents completing command-line tasks, scoring 65.4% on maximum effort settings. The Terminal-Bench project’s public leaderboard shows a separate entry for Opus 4.6, with one configuration scoring 62.9%.
Regarding GDPval-AA, a benchmark that measures performance on professional tasks across finance, legal, and other fields, Anthropic said Opus 4.6 beats OpenAI’s GPT-5.2 by about 144 Elo points, a difference that equates to a win rate of about 70% in a direct comparison. Artificial Analysis, which manages the GDPval-AA leaderboard, describes its evaluation framework in its methodology document.
Anthropic also cited results from BrowseComp, an OpenAI benchmark for browsing agents that measures the ability to find hard-to-find information across 1,266 questions that require persistent web navigation.
