A radar over AI and agents. Every morning, and only what several independent sources are saying at once.
The day in briefwritten by the model
01Xiaomi MiMo-V2.6 arrives on OpenRouter and vLLMXiaomi's MiMo-V2.6 model series with a 1M-token context window launched simultaneously on both OpenRouter and the vLLM framework.→ 2023
02OpenAI and Anthropic launch new model familiesBoth OpenAI and Anthropic introduced new model lines accompanied by significantly reduced operational costs compared to their previous iterations.→ 0912
03Google tests TPU processors in space orbitGoogle deployed satellite prototypes carrying Tensor Processing Units to evaluate machine learning workloads under harsh space conditions.→ 0217
Video of the day
reel no. 1 · 1:11
LABELarchitecture · 22 voices
Anthropic launches Claude Opus 5.5 with 40% lower operational costs
Anthropic has introduced Claude Opus 5.5, the first model in the Claude 5.5 family, delivering performance comparable to Claude Fable 5.1 while costing 40% less to run than Opus 5. The model features improved clarity and token efficiency across all effort levels.
Claude Code officially released cloud sessions allowing tasks to run while laptops are closed, and added a graceful stopping point when hitting the 5-hour limit. The platform also resumed charging for requests blocked by safeguards.
Google's Project Suncatcher is testing a TPU-powered satellite prototype aboard a SpaceX mission. Additionally, Google DeepMind published essays addressing misbehavior control in agent swarms and equitable AGI benefit distribution.
Hugging Face is conducting an extensive review of AI agents' internet access during training and evaluation, publishing ongoing transparency summaries. Meanwhile, new diarization models were released to track multiple speakers.
ChatGPT Voice added support for plugins like email, calendar, and Slack across web and mobile platforms. The ecosystem also prepares tiered subscription plans scaling up to a $500 Pro Max tier.
Google launched Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, featuring over 2,000 production-ready voices, support for 100 languages, voice design tools, and voice replication.
Microsoft has announced a major update to GitHub Copilot, introducing proactive Autopilot and two new models: GPT-6 Sol for interactive coding and lightweight GPT-6 Luna. The release also includes a new Copilot application featuring Home, Code, and Copilot modes.
DigitalOcean has released Managed Agents in public preview, enabling users to run Claude Code, Codex, or LangGraph agents in runtime environments that pause when idle. The service places agent tools behind a governed endpoint and supports over 75 open-source options.
Grok 4.7 xHigh achieved 58% on the Artificial Analysis AA-Briefcase benchmark, ranking just one point behind Claude Fable 5.1 Max at 59%. Additionally, the stealth model Pixel Canary launched on Cline, tying with GPT-6 Astra and outperforming Kimi K3 on the Next.js Agent Evals benchmark.
OpenAI has launched GPT-6 Sol and GPT-6 Luna, bringing the core strengths of GPT-6 Astra into faster and more affordable models designed for scaled workloads. Both models launch with API prices 50% lower than GPT-5.6, accompanied by Tesseract, a video creative suite built for AI agents.
At Meta Connect, Meta announced strategic integration partnerships for its personal AI agent Muse with Shopify, Expedia, and PayPal. These collaborations enable users to streamline shopping, trip planning, and checkout processes through Muse across participating merchants worldwide.
Anthropic has launched a new portal allowing developers to submit plugins, track reviews, and monitor usage for Claude. These plugins package the Model Context Protocol (MCP) along with custom skills, serving as the standard mechanism for building applications on Claude.
Anthropic has introduced Claude Opus 5.5, the first model in the Claude 5.5 family, delivering performance comparable to Claude Fable 5.1 while costing 40% less to run than Opus 5. The model features improved clarity and token efficiency across all effort levels.
Claude Sonnet 5.5 is currently undergoing stealth testing featuring a 1-million-token context window and a 128K-token maximum output length. Pricing is set at $2 per 1M input tokens and $10 for output tokens, arriving as a direct response to rival model releases.
Anthropic has increased the performance speed of claude.ai and its desktop application by 3x within two weeks. The development team shared specific prompts and techniques utilizing Claude to measure, debug, and improve execution speed.
SpaceXAI has released Grok 4.7, described as its most capable model yet for coding and knowledge work. The development of this system was supported by NVIDIA accelerated computing infrastructure.
Jürgen Schmidhuber, a pioneer in meta-learning, recursive self-improvement, and world models since the 1990s, has joined Sakana AI as Chief Scientific Advisor. Concurrently, the company introduced Agora-2, a next-generation multi-agent world model supporting real-time interaction for up to 20 humans and agents.
Google has partnered with Planet to launch a prototype satellite carrying four Tensor Processing Units (TPUs) into orbit. The mission evaluates how the hardware withstands harsh space conditions to assess the feasibility of hosting machine learning in orbit.
Apple has released a model on Hugging Face built by finetuning Qwen3.5-9B to convert long documents into small page images to conserve tokens, subsequently retrieving full text only for the pages relevant to a user's query.
The agent building community is experimenting with System One models like Jev and the Contrastive Language Model within custom harnesses. These models serve as high-speed routers or verifiers to improve performance for agent tasks.
OpenRouter added several new models, including Xiaomi's multimodal MiMo-V2.6 series with a 1M-token context window and the System One model Jev. Developers are testing these models for applications requiring fast response times.
MiMo-V2.6-Pro achieved the top spot on the Artificial Analysis Intelligence Index for open weights models. The model utilizes a classic Grouped Query Attention architecture and delivers competitive cost efficiency per task.
Researchers released a post-training approach that helps agents learn from real user sessions by imitating successful trajectories and correcting errors. Another study from Google evaluates the impact of distilling agent harnesses on task success rates.
The vLLM framework deployed day-one support for the Xiaomi MiMo-V2.6 models and DiffusionGemma-Jev. Both the Pro and Flash variants provide long-context and multimodal processing capabilities on the serving platform.
Tencent ARC Lab introduced the GameHorizon Suite, a data and evaluation framework designed to measure AAA gameplay capabilities across multiple temporal horizons for vision-language models, GUI agents, and coding agents.
The winner of Anthropic's hackathon has open-sourced their complete Claude Code setup, including agent skills, plugins, and practical tips. Concurrently, new research introduces methods to evolve agent skills while achieving 40 to 70 percent lower token costs compared to frontier methods, alongside NVIDIA's work on compiling public agent skills into reinforcement learning environments.
Alibaba aims to reach 20 gigawatts of global data center capacity by 2032 to support its artificial superintelligence ambitions, alongside plans for a 5 to 10 trillion parameter Qwen model and proprietary AI chips. Additionally, the Qwen team released RecreationWorld on Hugging Face as a sandbox for hybrid computer-use agents to explore, implement, and visually verify builds.
QUESTION — How can vision-language models (VLMs) be integrated into a visual action workspace to pilot robot manipulation more effectively?
On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone.
QUESTION — How can diverse edge-case tasks be automatically generated to evaluate and optimize tool-calling agents without human annotation?
Finetuning on data generated by EdgeGen yields a consistent mean progress improvement of 2 percent to 42 percent on tau2bench airline domain, while other baseline methods show degradation for some models.
QUESTION — How can we simultaneously improve reasoning accuracy and inference efficiency in language models by combining independently learned capabilities?
On Qwen3.5-4B, it raises HMMT 2025 accuracy from 59.2% to 64.0% with 10.7% fewer response tokens, and LiveCodeBench v5 accuracy from 41.7% to 54.2% with 9.6% fewer response tokens.
QUESTION — How can hidden chain-of-thought traces be extracted and characterized from closed-source frontier language models via standard APIs?
This research leverages a simple custom tool registered through a standard API feature to induce frontier models to externalize intermediate reasoning that is otherwise hidden in closed systems. By comparing against native CoT on open models and extending to systems like GPT-6 Astra, the authors demonstrate that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines across competition mathematics, science, and code generation. Structural analysis reveals that models like Astra exhibit token-efficient directed reasoning, resolving elementary steps internally while externalizing only crucial reasoning.
QUESTION — Can the resource-intensive teacher-critic stack in video generation post-training be completely eliminated?
With only 20 generator updates, the recipe reaches 84.87 on the official VBench evaluation, outperforming the previous best few-step causal baseline by 0.36, while requiring 16 A100 GPU-hours.
QUESTION — How can alignment failures in deployed language models be detected efficiently and cost-effectively without relying on expensive generative judges?
A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks.