Not another AI news riverOfficial evidence firstOne event can belong to many entitiesAction required is separate from popularityCommunity discussion never substitutes for fact
DeepSeek-V4-Flash Update The official release of the DeepSeek-V4-Flash API is now in public beta. The API calling method remains unchanged — simply set the model name to deepseek-v4-flash to use the latest version. Significantly enhanced agent capabilities, with benchmark results far exceeding V4-Pro-Preview: Terminal Bench 2.1: 82.7 NL2Repo: 54.2 Cybergym: 76.7 DeepSWE: 54.4 Toolathlon verified: 70.3 Agent Last Exam: 25.2 Automation Bench (Public): 25.1 DSBench-FullStack: 68.7 DSBench-Hard: 59.6 Note 1: For the Code Agent tasks in the public benchmark sets, the official DeepSeek-V4-Flash was tested using the DeepSeek Harness minimal mode (to be released soon) as the framework, with the max effort level, topp=0.95, and temperature=1.0 Note 2: DSBench-FullStack is an internal full-stack development test set, and DSBench-Hard is an internal Coding Agent hard-problem test set The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex. For the specific configuration, please refer to the documentation . DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained. Note: This upd
Action requiredReview affected integrations and migrate before the documented retirement boundary.
Grok 4.6 Grok 4.6, SpaceXAI's frontier model for coding, agentic tasks, and knowledge work, is now available on the xAI API. It has a 500k context window, text and image inputs with text-only output, and no text output limit. Pricing is $2 / $0.50 / $6 per 1M tokens (input / cached input / output) below 200k prompt tokens, and $4 / $1 / $12 above. Reasoning effort supports low, medium, high (default), and xhigh. See the Grok 4.6 overview and the announcement . August 11
Grok 4.5 Grok 4.5, SpaceXAI's model for coding, agentic tasks, and knowledge work, is now available on the xAI API. Priced at $2 / 1M input tokens and $6 / 1M output tokens, with configurable reasoning effort (low, medium, or high; default high). See the Grok 4.5 overview and the announcement .
October 15, 2025 Claude Haiku 4.5 launch We released our fastest, most cost-efficient model – Claude Haiku 4.5. Our latest small model matches Sonnet 4’s performance on coding, computer use, and agent tasks. Claude in Chrome updates Powered by Haiku 4.5: Claude in Chrome now defaults to Haiku 4.5 so it’s a faster, more responsive experience. You can always switch back to Sonnet 4.5. Claude handles image uploads for you: Give Claude an image and tell it where to upload, whether it’s an expense report, form attachment, or a picture upload. Show Claude exactly what you mean: Take a screenshot or drag to highlight specific parts of your screen. Point Claude to the exact button, field, or detail—much faster than describing complex layouts in words.
deepseek-chat The deepseek-chat model has been upgraded to DeepSeek-V2-0628. Model's reasoning capabilities have improved, as shown in relevant benchmarks: Coding: HumanEval Pass@1 79.88% -> 84.76% Mathematics: MATH ACC@1 55.02% -> 71.02% Reasoning: BBH 78.56% -> 83.40% In the Arena-Hard evaluation, the win rate against GPT-4-0314 increased from 41.6% to 68.3%. The model's role-playing capabilities have significantly enhanced, allowing it to act as different characters as requested during conversations.
Grok Build Grok Build is now available in beta. Use the interactive TUI, run headlessly in scripts, or build apps and orchestrators with the Agent Client Protocol. Install with a single command: Bash curl -fsSL https://x.ai/cli/install.sh | bash For more details, see the Grok Build docs . May 1
September 1, 2026 Claude Fable 5.1 and Claude Mythos 5.1 launch We just launched Claude Fable 5.1 and Claude Mythos 5.1, the world’s most advanced models for coding and knowledge work. For more information, see our blog post: Claude Fable 5.1 and Mythos 5.1 .
June 30, 2026 Claude Sonnet 5 launch We launched Claude Sonnet 5, our most agentic Sonnet model yet, with substantial improvements over Sonnet 4.6 in reasoning, tool use, coding, and knowledge work. For more information, see our blog post: Introducing Claude Sonnet 5 .
May 28, 2026 Claude Opus 4.8 launch We’ve upgraded Claude Opus to a new version. Claude Opus 4.8 shows improvements over Opus 4.7 in coding, agentic skills, reasoning, and practical knowledge work tasks. For more information, see our blog post: Introducing Claude Opus 4.8 . Enterprise plans can manage connector access with custom roles We added connector permissions to extend the existing custom roles framework and allow administrators to control which connectors, and which individual tools on those connectors, are available to each custom role. For more information, see Manage custom roles on Enterprise plans .
April 16, 2026 Claude Opus 4.7 launch Our latest model, Claude Opus 4.7, is now generally available. Opus 4.7 shows improvements in software engineering and complex, long-running coding tasks, as well as better vision, allowing it to see images in higher resolution. For more information, see our blog post: Introducing Claude Opus 4.7 .
February 17, 2026 Claude Sonnet 4.6 launch We launched our most capable Sonnet model yet, with a full upgrade of the model’s skills across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. Sonnet 4.6 also features a 1M token context window in beta. Read our blog post for more information: Introducing Claude Sonnet 4.6 .
February 5, 2026 Claude Opus 4.6 launch We’ve upgraded our smartest model and improved its coding skills. Read our blog post for more information: Introducing Claude Opus 4.6 . Introducing Claude for PowerPoint Claude is now available as an add-in for PowerPoint. Read more here: Use Claude for PowerPoint . Claude for Excel improvements We’ve updated Claude for Excel so it uses Opus 4.6 and supports native Excel operations such as pivot table editing and conditional formatting. See our updated article for more information: Using Claude for Excel .
January 12, 2026 Cowork research preview on Claude Desktop (macOS only) for Max plans Cowork brings Claude Code's agentic capabilities to the Claude desktop app for knowledge work beyond coding. It runs locally on your computer in an isolated VM, enabling direct access to local files and MCP integrations. Refer to this article to learn more: Getting started with Cowork . Health and fitness data on Claude Mobile Claude can now read and analyze your health and fitness data on iOS and Android. Ask Claude about your activity patterns, workout trends, sleep quality, and more—Claude will provide insights and visualizations using native charts. Health features are available on Pro and Max plans and currently limited to users in the US. On Android, Health Connect and Android 14 or later are required. See the following articles for more information: Using Claude with iOS Apps Using Claude with Android Apps HIPAA-ready Enterprise plans We now offer a HIPAA-ready version of Claude that is available for organizations with Enterprise plans that choose to process protected health information (PHI) through Claude. See HIPAA-ready Enterprise plans for more information.
September 29, 2025 Claude Sonnet 4.5 launch We released our newest model, Sonnet 4.5. This is the best model in the world for real-world agents, coding, and computer use. Read our blog post here: Claude Sonnet 4.5 . Creating and editing files with Claude for Pro plans and mobile Pro users can now leverage Claude’s file creation and editing capabilities, and users on all paid plans can access these features on Claude for iOS or Android. See this updated article for more information: Create and edit files with Claude . Claude in Chrome updates The remaining Max users on our waitlist were granted access to Claude in Chrome, along with the following updates: Powered by Sonnet 4.5: Claude in Chrome now defaults to Sonnet 4.5, our smartest model yet. Improved for browser tasks, you'll notice better reasoning, fewer errors, and more reliable task completion—especially for multi-step workflows. Work across multiple tabs: Claude can now juggle multiple browser tabs at once. Just drag tabs into Claude's tab group and it can see and work across all of them simultaneously—no more jumping back and forth to gather information before taking action. Smarter on the sites you use every day: Claude n
deepseek-coder The deepseek-coder model has been upgraded to DeepSeek-Coder-V2-0614, significantly enhancing its coding capabilities. It has reached the level of GPT-4-Turbo-0409 in code generation, code understanding, code debugging, and code completion. Additionally, it possesses excellent mathematical and reasoning abilities, and its general capabilities are on par with DeepSeek-V2-0517.
Confidence, importance, and discussion are three different numbers.
Official and first-party repository changes may publish automatically after deterministic validation. Independent sources remain draft until review. Community links are attached as discussion evidence and cannot create an indexable fact by themselves.
Follow the systems that can break your product or budget.
Save events and hubs locally today. Confirm your email for APIDir research updates, or request team access for entity watchlists, migration alerts, API exports, and webhooks.