Claude Code has officially launched Cloud sessions out of research preview, allowing tasks to continue running even when laptops are closed. Existing subscribers receive one-time trial credits of $100 on Pro and $250 on Max.
OpenAI has rolled out the GPT-6 Sol and Luna models across ChatGPT Work, Codex, Plus, Pro, Business, Enterprise, and Edu tiers. Both models are also made available through the API.
Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its new expressive audio generation models supporting 100 languages. These models feature over 2,000 production-ready voices along with voice design and replication capabilities.
ChatGPT Voice has been updated to integrate plugins such as email, calendar, and Slack. The voice capability is powered by GPT-6 Astra, Sol, and Luna across ChatGPT Work on web and mobile.
Hugging Face has released the Nemotron 3 Diarization model featuring 100 million parameters to track up to eight speakers during overlapping conversations. The organization is also conducting an extensive review regarding agents' internet access during training and evaluation.
After Effects and Premiere have been rebuilt as Tesseract to provide direct access for GPT-6 Astra and Sol AI agents. The platform connects agents directly to the creative engine to execute design and video workflows.
Microsoft rolled out a major update to GitHub Copilot, introducing the proactive Autopilot feature alongside two new models: GPT-6 Sol for interactive and agentic coding, and the cost-effective GPT-6 Luna.
Anthropic launched Claude Opus 5.5, the first model in its new 5.5 family. It matches the performance of the Fable 5.1 level while reducing running costs by 40% compared to Opus 5.
Personal AI agent platform Muse announced partnerships with Shopify, Expedia, and PayPal to handle shopping, travel, and checkout workflows. The company also introduced Muse Charm, a keychain device scheduled to ship in December.
SpaceXAI released Grok 4.7, scoring 46 on the Artificial Analysis Intelligence Index and outperforming GPT-5.6 Sol on the Coding Agent Index. The release places SpaceXAI third among major AI labs for agentic coding.
Sakana AI announced Agora-2, a multi-agent world model supporting up to 20 humans and agents interacting in a real-time shared simulation. Additionally, Jürgen Schmidhuber joined the organization as Chief Scientific Advisor.
Grok 4.7 xHigh reached 58% on Artificial Analysis' AA-Briefcase benchmark, placing just one point behind Claude Fable 5.1 Max. Additionally, a stealth model named Pixel Canary launched with performance matching GPT-6 Astra on Next.js agent evaluations.
Apple released a Qwen3.5-9B finetune on Hugging Face designed to convert long documents into page images to save tokens. Other releases include the tev1-4B-experimental classifier and PrismML Bonsai 2 27B, focusing on smaller footprints and local hardware compatibility.
New research demonstrates that implementing an ontology layer for AI agents significantly enhances their capabilities. The study shows that GPT-5.5 gains 26.7 points on DDR-Bench when utilizing self-evolving ontologies.
DeepSeek is currently training a 2-trillion-parameter model, with an 8-trillion-parameter model also planned. Additionally, the company expects to receive new Huawei training chips between Q4 2026 and Q1 2027.
Google has officially introduced a new laptop product line named Googlebook, joining its existing hardware categories to address modern workplace computing needs.
OpenRouter has expanded its platform with new model integrations, including Space Bunny Alpha, Typesafe AI's high-speed decision model Jev, and Xiaomi's MiMo-V2.6 series. These additions feature multi-modal capabilities across text, image, video, and audio, alongside 1-million-token context windows.
Xiaomi has introduced MiMo-V2.6-Pro, which debuts as the top open-weights model on the Artificial Analysis Intelligence Index at a cost of $0.13 per task. Built on a classic Grouped Query Attention architecture, the model achieves leading performance in weighted average benchmarks.
The release of System One models like Jev and Contrastive Language Model (CLM) has spurred the development of custom agent harnesses. New optimization tools, such as ReAnchor built with DSPy methodology, leverage these models to enhance verification and execution speed.
Google DeepMind has released three new institute essays examining how to control misbehavior in agent swarms, orchestrate complex networks of humans and AIs, and ensure the equitable distribution of AGI benefits.
vLLM has rolled out day-zero support for Xiaomi's MiMo-V2.6 model family, spanning a 1.02T total parameter Pro version and a 309B parameter Flash checkpoint. Additionally, disaggregated setups achieved 469 tok/s single-user decode performance using vLLM on AMD hardware.
Recent research introduces improved methods for evolving agent skills by ranking candidates with a language model, achieving 40% to 70% lower token costs. Work from NVIDIA, such as Skill2Env, further demonstrates compiling public agent skills into reinforcement learning environments.
Alibaba has targeted 20 gigawatts of global data center capacity by 2032 alongside plans for a 5–10 trillion parameter Qwen model and proprietary AI chips. Simultaneously, the Multilingual TTS Arena leaderboards debuted to compare text-to-speech models across nine languages.
QUESTION — Do modern embodied agents actually follow language instructions, or do they exploit low scene entropy?
Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup.
QUESTION — How does the Mixture-of-Experts architecture of Hunyuan-A13B optimize inference performance and computational efficiency?
It contains 80 billion total parameters but activates only 13 billion during inference, balancing model capability, computational efficiency, and deployment cost.
QUESTION — Do video generation models possess inherent object permanence, and can training them on a core-cognition inspired dataset instill this physical intelligence?
WROP comprises 150 hand-designed tasks divided into six cognitive categories.
QUESTION — How are underlying physical laws and constraints embedded into machine learning algorithms for robotics?
We adopt a unified taxonomy that classifies existing approaches according to their physics embedding: physics-guided inputs, data, and representations; physics-encoded model architectures; and physics-informed training loss functions.
QUESTION — How can we build a computer vision architecture that achieves high few-shot accuracy under extremely constrained parameter and sample budgets?
The architecture contains 22,249-34,917 parameters.
QUESTION — How can a large-scale dataset with high-fidelity annotations be constructed to improve RGB-D semantic segmentation?
RGBD20K offers 20,000 RGB-D image pairs, providing a substantially larger training resource that benefits the development of more powerful deep models.