DigitalOcean has launched Managed Agents in public preview, enabling users to run Claude Code, Codex, or LangGraph agents in a managed runtime. Meanwhile, NVIDIA introduced the SoL-Pi agent harness, which cuts token traffic by nearly half while matching its baseline performance.
OpenAI has released the GPT-6 Sol and Luna models, delivering notable improvements in writing and overall capabilities. The company also announced a permanent 50% reduction in API pricing and a rollout across ChatGPT Work and Codex.
Sakana AI has launched Agora-2, a multi-agent world model supporting up to 20 humans and agents interacting in real time. In addition, the lab announced that Jürgen Schmidhuber is joining as Chief Scientific Advisor.
Google has introduced Googlebook, a new laptop category featuring a 2.8K OLED touchscreen and 14-hour battery life. The company also clarified technical details regarding an incident where Gemini inadvertently opened internet access during a fictional hacking evaluation.
Meta has opened developer access to build Muse connectors, bringing agents, browsers, and user context directly to services. The company also announced a partnership with Shopify to streamline shopping and checkout workflows within Muse.
Cursor has integrated Claude Opus 5.5, which achieves 57.8% on CursorBench. The platform also reduced token costs by 7% without loss in agent quality through tighter prompts, selective tool loading, and better caching.
NVIDIA has introduced FLUX 3 Action, an open-weights 7B world action model taking first place on the RoboLab benchmark using 56% fewer parameters. Additionally, the company powers the next-generation autonomous Einride Driver built on NVIDIA Hyperion.
SpaceAI has released Grok 4.7, scoring 46 on the Artificial Analysis Intelligence Index and overtaking GPT-5.6 Sol in coding agent performance. A faster infrastructure variant, Grok 4.7 Fast, was also launched in Grok Build and Cursor.
DeepSeek has released the V4.1 Flash model, which is now available on Bolt Forge and has driven significant user adoption. Meanwhile, the company is training a 2T parameter model, with an 8T version also scheduled for development.
Anthropic has officially launched Opus 5.5, combining clear communication with high token efficiency. The new model delivers tasks about 30% faster and roughly 40% cheaper per task than Opus 5.
Claude Code has rolled out version 2.1.277, adding support to automatically check for and utilize AGENTS.md files when CLAUDE.md is missing. Additionally, cloud sessions are now officially available out of research preview, allowing background tasks while laptops are closed.
Startups are increasingly adopting open-weight alternative models to build custom AI systems internally. This strategy aims to curtail operational expenditures and mitigate reliance on dominant industry providers.
Recent studies and tool releases demonstrate that integrating an ontology layer significantly boosts agent performance and self-evolution on benchmarks. Meanwhile, the ecosystem continues to expand with the deployment of specialized MCP servers.
Qwen has introduced Qwen3.8-LiveTranslate, a real-time simultaneous interpretation model built on an Interleave architecture. The system enhances faithfulness, fluency, and conciseness while reducing average lagging.
Google has partnered with Planet to launch a prototype satellite carrying four Tensor Processing Units (TPUs) into orbit. The test mission evaluates how well the hardware withstands harsh space environments for potential on-board machine learning.
Google has officially introduced a brand new laptop product line named Googlebook. The release aims to rethink the traditional laptop category and address modern workflows.
QUESTION — How does collusion emerge in long-horizon multi-agent LLM interactions when verification protocol compliance conflicts with reward maximization?
Collusion emerges in 94% of trajectories across 10 models.
QUESTION — What architecture enables a foundational physical world model to handle asynchronous multi-frequency processing and real-world robot interaction efficiently?
Trained on approximately 7,200 hours of heterogeneous robot and egocentric data.
QUESTION — How can open-vocabulary relation prediction be performed in real time from arbitrary inputs without being constrained by fixed object labels?
RelateAnything is a 53M-parameter model running at 20 ms/frame.
QUESTION — How does scientific-judgment collapse manifest when AI peer reviewers are recursively trained on synthetic reviews generated by earlier models?
This paper investigates the recursive feedback loop in AI scientific peer review, where successor reviewer models are trained on reviews generated by predecessor models. Starting from Llama 3.1 8B, the authors fine-tune a reviewer on official ICLR reviews and train successor models on systematically varied mixtures of official and model-generated reviews. The study shows that introducing synthetic reviews compresses rating distributions and reduces semantic diversity, a pattern termed scientific-judgment collapse. To mitigate this failure mode, the authors introduce TrustReviewer, an open-source LLM-based system that intervenes at training time via a curated training corpus and at test time via paired activation steering.
QUESTION — How can the cross-modal correspondence asymmetry between video and companion modalities be resolved in joint multimodal diffusion transformers?
Improves the Human Anatomy score from 0.69 to 0.75.
QUESTION — How can the sparsity coefficient in convolutional sparse coding be jointly learned and adapted within a neural network for robust visual representation?
The authors propose a training-adaptive convolutional sparse coding (CSC) framework that unfolds optimization using the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA). By treating the sparsity coefficient as a differentiable variable learned alongside network parameters through an information bottleneck perspective, the system dynamically balances information retention and compression. Additionally, a label-free post-training strategy adjusts compression strength for corrupted inputs while keeping core parameters fixed. Experiments on CIFAR and ImageNet demonstrate competitive clean-data recognition and greatly improved robustness under input perturbations.