Not another AI news riverOfficial evidence firstOne event can belong to many entitiesAction required is separate from popularityCommunity discussion never substitutes for fact
DeepSeek-V4-Flash Update The official release of the DeepSeek-V4-Flash API is now in public beta. The API calling method remains unchanged — simply set the model name to deepseek-v4-flash to use the latest version. Significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview: Terminal Bench 2.1: 82.7 NL2Repo: 54.2 Cybergym: 76.7 DeepSWE: 54.4 Toolathlon verified: 70.3 Agent Last Exam: 25.2 Automation Bench (Public): 25.1 DSBench-FullStack: 68.7 DSBench-Hard: 59.6 Note 1: For the Code Agent tasks in the public benchmark sets, the official DeepSeek-V4-Flash was tested using the DeepSeek Harness minimal mode (to be released soon) as the framework, with the max effort level, topp=0.95, and temperature=1.0 Note 2: DSBench-FullStack is an internal full-stack development test set, and DSBench-Hard is an internal Coding Agent hard-problem test set The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex. For the specific configuration, please refer to the documentation . DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained. Note: This upd
Action requiredReview affected integrations and migrate before the documented retirement boundary.
Grok 4.6 Grok 4.6, SpaceXAI's frontier model for coding, agentic tasks, and knowledge work, is now available on the xAI API. It has a 500k context window, text and image inputs with text-only output, and no text output limit. Pricing is $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200k prompt tokens, and $4 / $1 / $12 above. Reasoning effort supports low, medium, high (default), and xhigh. See the Grok 4.6 overview and the announcement . August 11
Grok 4.5 Grok 4.5, SpaceXAI's model for coding, agentic tasks, and knowledge work, is now available on the xAI API. Priced at $2 / 1M input tokens and $6 / 1M output tokens, with configurable reasoning effort (low, medium, or high; default high). See the Grok 4.5 overview and the announcement .
deepseek-reasoner deepseek-reasoner is our new model DeepSeek-R1. You can invoke DeepSeek-V3 by specifying model='deepseek-reasoner' . For details, please refer to: DeepSeek-R1 Release For guides, please refer to: Thinking Mode
deepseek-chat The deepseek-chat model has been upgraded to DeepSeek-V2-0628. Model's reasoning capabilities have improved, as shown in relevant benchmarks: Coding: HumanEval Pass@1 79.88% -> 84.76% Mathematics: MATH ACC@1 55.02% -> 71.02% Reasoning: BBH 78.56% -> 83.40% In the Arena-Hard evaluation, the win rate against GPT-4-0314 increased from 41.6% to 68.3%. The model's role-playing capabilities have significantly enhanced, allowing it to act as different characters as requested during conversations.
June 30, 2026 Claude Sonnet 5 launch We launched Claude Sonnet 5, our most agentic Sonnet model yet, with substantial improvements over Sonnet 4.6 in reasoning, tool use, coding, and knowledge work. For more information, see our blog post: Introducing Claude Sonnet 5 .
May 28, 2026 Claude Opus 4.8 launch We’ve upgraded Claude Opus to a new version. Claude Opus 4.8 shows improvements over Opus 4.7 in coding, agentic skills, reasoning, and practical knowledge work tasks. For more information, see our blog post: Introducing Claude Opus 4.8 . Enterprise plans can manage connector access with custom roles We added connector permissions to extend the existing custom roles framework and allow administrators to control which connectors, and which individual tools on those connectors, are available to each custom role. For more information, see Manage custom roles on Enterprise plans .
February 17, 2026 Claude Sonnet 4.6 launch We launched our most capable Sonnet model yet, with a full upgrade of the model’s skills across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. Sonnet 4.6 also features a 1M token context window in beta. Read our blog post for more information: Introducing Claude Sonnet 4.6 .
September 29, 2025 Claude Sonnet 4.5 launch We released our newest model, Sonnet 4.5. This is the best model in the world for real-world agents, coding, and computer use. Read our blog post here: Claude Sonnet 4.5 . Creating and editing files with Claude for Pro plans and mobile Pro users can now leverage Claude’s file creation and editing capabilities, and users on all paid plans can access these features on Claude for iOS or Android. See this updated article for more information: Create and edit files with Claude . Claude in Chrome updates The remaining Max users on our waitlist were granted access to Claude in Chrome, along with the following updates: Powered by Sonnet 4.5: Claude in Chrome now defaults to Sonnet 4.5, our smartest model yet. Improved for browser tasks, you'll notice better reasoning, fewer errors, and more reliable task completion—especially for multi-step workflows. Work across multiple tabs: Claude can now juggle multiple browser tabs at once. Just drag tabs into Claude's tab group and it can see and work across all of them simultaneously—no more jumping back and forth to gather information before taking action. Smarter on the sites you use every day: Claude n
DeepSeek-V3.1 Both deepseek-chat and deepseek-reasoner have been upgraded to DeepSeek-V3.1. deepseek-chat corresponds to DeepSeek-V3.1's non-thinking mode , while deepseek-reasoner corresponds to its thinking mode . Key updates in DeepSeek-V3.1: Hybrid reasoning architecture : A single model supports both thinking mode and non-thinking mode Improved reasoning efficiency : Compared to DeepSeek-R1-0528, DeepSeek-V3.1-Think provides answers in significantly less time Enhanced agent capabilities : With post-training optimization, the new model achieves major improvements in tool usage and intelligent agent tasks SWE-bench Verified: 66.0 SWE-bench Multilingual: 54.5 Terminal-bench: 31.3
deepseek-coder The deepseek-coder model has been upgraded to DeepSeek-Coder-V2-0614, significantly enhancing its coding capabilities. It has reached the level of GPT-4-Turbo-0409 in code generation, code understanding, code debugging, and code completion. Additionally, it possesses excellent mathematical and reasoning abilities, and its general capabilities are on par with DeepSeek-V2-0517.
Confidence, importance, and discussion are three different numbers.
Official and first-party repository changes may publish automatically after deterministic validation. Independent sources remain draft until review. Community links are attached as discussion evidence and cannot create an indexable fact by themselves.
Follow the systems that can break your product or budget.
Save events and hubs locally today. Confirm your email for APIDir research updates, or request team access for entity watchlists, migration alerts, API exports, and webhooks.