OpenAI has introduced the GPT-6 Sol and GPT-6 Luna models, building on the capabilities of GPT-6 Astra to offer faster and more affordable options. Additionally, the company launched Tesseract, a video creative suite designed to give AI agents direct access to video editing tools.
SpaceXAI has released Grok 4.7, which scores 46 on the Artificial Analysis Intelligence Index and places the lab in the top four AI labs. The model demonstrates improved coding agent performance while maintaining high speed and low cost.
Anthropic has introduced Claude Opus 5.5, the first model in the Claude 5.5 family, delivering performance comparable to Claude Fable 5.1 while cutting operating costs by 40% compared to Opus 5.
DigitalOcean has launched a public preview of Managed Agents, allowing users to run tools like Claude Code or Codex within runtime environments that pause automatically during idle periods.
NVIDIA's accelerated computing infrastructure is supporting models across the AI ecosystem, including powering SpaceXAI's Grok 4.7 for demanding coding and knowledge workflows.
OpenAI updated ChatGPT voice capabilities to integrate plugins across email, calendar, and Slack. The platform also added support for linking multiple accounts with plugins and running Chrome extensions inside the desktop app.
Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS featuring over 2,000 production-ready voices and support for 100 languages. The release also extends real-time audio capabilities through Gemini 3.8 Live in search.
DeepSeek is currently training 2-trillion and 8-trillion parameter models, with plans to incorporate Huawei training chips in late 2026 or early 2027. The platform also integrated the DeepSeek V4.1 Flash model into developer tools.
The Muse app for Mac has been released, operating across files, calendar, notes, and messages. The platform also opened API connector access for developers and partnered with Shopify to streamline shopping workflows.
Google launched Googlebook, a new laptop featuring a 2.8K OLED touchscreen display and a 14-hour battery life. The device is built on the Android technology stack paired with ChromeOS desktop foundations.
Anthropic partnered with Accenture for independent evaluations of frontier AI systems, committing at least $1 billion to building capacity. The company has also established a physical wet lab in San Francisco for biological research.
NVIDIA released the Nemotron 3 Diarization model with 100 million parameters, capable of tracking up to eight speakers and handling overlapping voices. The system identifies who spoke and when within multi-person audio transcripts.
Claude Code updated to version 2.1.277, adding support for AGENTS.md configuration files when CLAUDE.md is absent. Cloud sessions also exited research preview, enabling Claude to run parallel threads even after laptops are closed.
Google vừa công bố CC, một trợ lý AI được thiết kế để hỗ trợ các gia đình quản lý lịch trình và công việc chung. Trong một diễn biến khác, các lãnh đạo công nghệ lớn được cho là đã thúc ép chính quyền loại bỏ một đề xuất quản lý AI do người đứng đầu Google DeepMind đề xuất.
Recent research demonstrates that incorporating a self-evolving ontology layer significantly improves agent performance on complex benchmarks. Meanwhile, the ecosystem continues to expand with new integrations and servers built around the Model Context Protocol.
PixVerse has released R2, a real-time world model that allows users to interact with and edit dynamic environments using prompts. Additionally, recent studies highlight that employing an explicit graph harness can drastically improve a model's performance on long-horizon robotic tasks without altering the underlying weights.
Ternary Bonsai 2 27B has been released, utilizing a 9x reduction in size compared to its Qwen3.8 base model while retaining 98.2% of its aggregate benchmark performance. Additionally, the Qwen3.8 ecosystem expanded with LiveTranslate, a simultaneous interpretation model built on an interleave architecture.
MiMo-V2.6-Pro has debuted as the top open weights model on the Artificial Analysis Intelligence Index. Built on a classic grouped query attention architecture, it establishes a strong position on the intelligence versus cost per task Pareto frontier.
Grok 4.7 xHigh has climbed close to the top spot on Artificial Analysis' AA-Briefcase benchmark, trailing the leader by a single point. Meanwhile, indicators point toward an upcoming Kimi K3.1 release, alongside demonstrations of running large models locally on standard CPUs.
The community discussed the potential of automating scientific research using multi-agent systems. Conversations centered on the prospects and implications of superintelligence driving autonomous breakthroughs in mathematics and engineering.
Zhipu reported that GLM-5.3-Flash runs across more than 10,000 domestic AI accelerators, achieving a 3.2x increase in throughput through automated optimization driven by an AI agent. Meanwhile, models like DiffusionGemma and Ternary-Bonsai received updates within vLLM and MLX frameworks.
DeepSeek operates production-scale units spanning around 160 nodes, serving about 3 million sandboxes per day. Simultaneously, Xiaomi released the MiMo-V2.6 series featuring Pro and Flash models, with the Pro variant scoring 46 points on Artificial Analysis.
Developers deployed a 27-billion-parameter ternary model at 1.72 bits per weight on a single graphics card with 12GB of VRAM. Additionally, updates brought llama.cpp's Metal kernels to transformers and added SGLang support for image generation models.
Alibaba targeted 20 gigawatts of global data center capacity by 2032 alongside plans for a 5 to 10 trillion parameter Qwen model series. The Qwen team also released RecreationWorld, a five-platform sandbox environment for computer-use agents.
Researchers introduced a method allowing AI agents to improve skill prompts by ranking candidates, achieving 40% to 70% lower token costs compared to conventional methods. NVIDIA also detailed an approach to compile public agent skills into reinforcement learning environments.
Artificial Analysis introduced safety refusal reporting in version 1.5 of its Coding Agent Index to help explain model behavioral differences. Data showed distinct variations in safety refusal rates among evaluated models during coding tasks.
A joint university study introduced a path for financial agents to learn from SEC filing errors by requiring all new behaviors to pass regression tests first. New adaptive retrieval frameworks were also released to better handle specific data queries.
QUESTION — How can a full-duplex interaction system be built to integrate continuous perception, conversational control, and asynchronous background task execution using 9B models?
Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%).
QUESTION — How can heterogeneous enterprise document formats (PDFs, Word, scans) be converted and chunked into retrieval-optimized Markdown while minimizing token costs and processing time?
On the 236-document, 795-page PDF subset of the RAG-Multi-Corpus benchmark, D-RAC converts and chunks the entire corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks.
QUESTION — How can memory curation for LLM agents be optimized by deferring the summarization process from write-time to read-time?
Across ALFWorld, WebShop, and τ^2-bench, JitMem improves over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.
QUESTION — How can token-level credit be mathematically defined and leveraged to improve actor-critic training in LLM post-training?
In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively.
QUESTION — How can long-term episodic memory be provided to Vision-Language-Action models without bloating the context or increasing inference latency?
MemBodied achieves 7.81times the mean success rate of a stateless policy and 2.98times of vanilla recurrent memory across five evaluated RMBench tasks.
QUESTION — How can a decision-only judge reduce the inference cost of LLM-as-a-judge while preserving evaluation accuracy?
JEV is within three percentage points of a state-of-the-art LLM judge on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee.
QUESTION — What causes training instability in full-pipeline FP8 reinforcement learning for LLMs, and how can it be resolved?
Compounded FP8 quantization noise distorts the importance ratio, pushing negative-advantage tokens outside the trust region and zeroing out their gradients.
QUESTION — How can physical world representations be distilled from a world model into compact Vision-Language-Action policies without adding runtime inference latency?
The deployed policy runs in 32 ms and 1.86 GB on a consumer RTX 5090, identical to the undistilled baseline.
QUESTION — How can reinforcement learning fine-tuning be effectively applied to long-horizon real-world manipulation tasks with minimal human intervention?
On bimanual YAM and single-arm Franka tasks, PARTS improves complete-task success from 32% to 61% and from 50% to 95%, respectively, using tens of minutes of real-world RL rollouts per task on average.
QUESTION — How can reinforcement learning policies be optimized for open-ended SVG code generation without relying on ground truth data or human preference labels?
Prompting a vision-language judge with a multi-axis rubric correlates with human judgments far better than scalar metrics.
QUESTION — What memory mechanisms preserve temporal persistence and historical information in autoregressive video generation under bounded context limits?
This paper presents a systematic and comprehensive review of memory mechanisms in autoregressive (AR) video generation. As generated sequences expand, models face strict constraints on context windows, storage, and compute, causing critical historical information to leave the active context. The authors formulate memory operationally as persistent historical information maintained across outer AR steps to influence future generation. The literature is organized through five complementary perspectives: Forms, Functions, Operations, Learning, and Evaluation, establishing a structured foundation for building reliable memory-conditioned video generation systems.
QUESTION — How can deformable assets for robot manipulation be automatically generated from text or a single image using an agentic system?
DeformSmith generates assets with better visual quality and physical plausibility than state-of-the-art baselines, including PhysGen3D, PhysGM, and PhysX-Omni.
QUESTION — How can MLLMs be developed, trained, and evaluated for closed-loop, long-horizon robotic bin packing?
PackLab-VLM outperforms conventional packing heuristics, traditional reinforcement learning methods, and general-purpose MLLMs across object sets and container configurations.
QUESTION — How can a geometry-native latent space be constructed to improve 3D consistency and visual quality in world generation?
Replacing the latent with GAE improves visual quality and independently measured 3D coherence: FVD falls by 12.7% and 23.1% on RealEstate10K and DL3DV.
Checkable assertions with their sources and their contradictions. Each stands on at least two independent sources, or a person read it first; the stamp says which.
TWO SOURCES · NOT REREAD
01 · 24 Sept 2026 · claude · 2 sources
Claude discovered a previously uncharacterized enzyme system with features reminiscent of CRISPR after AI agents searched through a DNA sequence database.
“We’ve set up a molecular biology lab at Anthropic and we’re announcing our first discovery! Claude discovered a new CRISPR-like enzyme. 950 agents spent 21 hours searching through a database of DNA sequences until one of the agents found something striking”