A radar over AI and agents. Every morning, and only what several independent sources are saying at once.
145 ACCOUNTS HEARD
The day in briefwritten by the model
01Anthropic introduces Sonnet 5.5Anthropic has released Sonnet 5.5, delivering increased processing speed and lower costs while introducing enhanced defenses against distillation attacks.→ 11
Video of the day
reel no. 9 · 1:26
LABELarchitecture · 23 voices
Google launches Gemini 4 Argon with a 1 million token output limit
Google announced Gemini 4 Argon, delivering performance in complex workflows across software engineering and cybersecurity. The model features an industry-leading 1 million token output limit and is initially rolling out to trusted testers.
Claude Code now enables users to customize its interface and behavior using TypeScript-based mods shipped inside plugins. The platform also introduces a built-in 'You should know' plugin to scan model outputs for critical information.
Several new artificial intelligence models achieved leading positions on technical benchmarks and cybersecurity indices. Google released Gemini 4 Argon matching GPT-6 Astra at a lower cost, while Grok 4.7 claimed the top spot on the AA Cyber Index and Claude Opus 5.5 improved performance on Drone-Bench.
ChatGPT integrated plugin recommendations directly into conversations and brought additional server capacity online for the GPT-6.1 Sol model to handle high demand across subscriptions and APIs. The platform also expanded developer and consumer integrations.
Google announced Gemini 4 Argon, delivering performance in complex workflows across software engineering and cybersecurity. The model features an industry-leading 1 million token output limit and is initially rolling out to trusted testers.
The Muse project introduced Muse Gadgets, featuring open-source ESP32 firmware and a Linux SDK for building hardware devices compatible with Muse. This rollout accompanies a continuous series of model releases spanning code, image, and video generation.
Google released Gemini 4 Argon at a lower task cost compared to competing models such as Astra and Opus, while supporting long multi-step reasoning problems through its 1 million token output limit.
OpenAI has introduced the GPT-6 model family, featuring the hardware-accelerated GPT-6 Astra and the cost-efficient GPT-6.1 Sol. These models deliver significantly faster speeds and enhanced multimodal processing capabilities across paid plans and the API.
ChatGPT now enables users to build and deploy Model Context Protocol servers directly within the platform. This update streamlines data connection workflows and the integration of plugin extensions.
Coding assistant Devin now allows users to draw from their ChatGPT Plus or Pro subscription quotas directly. Additionally, the platform has rolled out substantial price reductions across its service tiers while improving overall capabilities.
Anthropic has launched a dedicated website for developers building applications with Claude. The new platform provides engineering deep dives, API guides, and practical documentation from the core development teams.
Anthropic has released Claude Sonnet 5.5, delivering a 30% increase in processing speed and reduced costs compared to the previous version. The model now also powers the free tier of the service.
The GLM model family, including versions 5.3 and 5.3 Flash, is now available within the Cursor development environment. These open-weight models achieve leading scores on CursorBench 4.0.
OpenRouter has integrated new models onto its platform, including Liquid AI's d1 decision model and the Pareto 26.10 preview. Meanwhile, Claude Opus 5.5 has rapidly captured the highest share of spend and tokens among Anthropic models on the service.
OpenAI has launched Dots, an always-on AI agent system operating 24/7 with access to a dedicated browser and thousands of applications. The system is designed to autonomously handle complex tasks such as managing customer service communications.
OpenAI has released Dots, a competitor to existing bot assistants on the market. Its primary differentiators include continuous availability and integrated access to all previous memories and interactions within the ChatGPT ecosystem.
Alibaba has released Qwen-Image-2.1, featuring a 7-billion parameter visual generator that claims the number one position on open-weights image leaderboards. Additionally, Qwen3.8 models have been integrated into multi-step workflows and live decision-making platforms.
Nvidia has partnered with over 100 industry organizations to introduce the Open Agent Safety Platform, combining OpenShell and Sentry. The platform provides robust security boundaries and sandboxing tools for long-running artificial intelligence agents.
Indications point to the upcoming reveal of an always-on agent from OpenAI, provisionally named o or Aeon. The release follows significant reported productivity gains from internal testing of precursor assistant technology.
OpenAI has reopened Pro subscriptions at $200 while altering how usage limits are calculated. The adjustment effectively halves the previous usage allowance for subscribers under the updated framework.
OpenAI has recorded a campaign attempting to extract the hidden reasoning capabilities of its models. According to OpenAI, a core cluster of this activity is attributed to individuals associated with Kimi developer Moonshot AI.
Researchers have released new frameworks and papers focusing on video world models and embodied physics reasoning. These initiatives aim to bridge the gap between realistic video generation and accurate physical laws.
DeepSeek temporarily paused its free promotion for the V4.1-Flash model due to abnormally high levels of abuse. The development team is actively investigating the situation to implement mitigation measures.
Community channels have published comprehensive recaps covering all major announcements, product launches, and industry updates from recent events. These resources serve as central hubs for tracking ongoing developments.
Technical discussions focused on implementations of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO). Meanwhile, new video generation models entered public evaluation leaderboards.
3 independent accounts7 posts3,216 interactions
deepseek r1reinforcement learning with verifiable rewards
The vLLM ecosystem added day-0 support for large-scale open-weights models like IQuest-Q1 and released updated semantic routing tools. Maintainers continue to optimize KV cache coordination and prefill-decode topologies for large deployments.
Google DeepMind published a comprehensive survey mapping out theoretical frameworks of machine consciousness. Researchers also highlighted future projections regarding autonomous AI agent token consumption and scientific applications.
LangChain has released mcp-adapters version 2.0, adding support for the latest stateless Model Context Protocol version for TypeScript agents and interactive tools. Additionally, the platform updated its LangChain Academy curriculum with a focus on Deep Agents and introduced a single command deployment tool for Managed Deep Agents.
QUESTION — How can a full-duplex interaction system be built to integrate continuous perception, conversational control, and asynchronous background task execution using 9B models?
Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%).
QUESTION — How can existing general-purpose agents be upgraded with omni-native production capabilities across multiple modalities without costly foundation model updates?
Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval.
QUESTION — How to autonomously construct specialized harnesses, improve them through experience, and orchestrate them across domains instead of manually engineering a single domain-specific harness?
Raven is an open-source multi-agent ecosystem that automatically constructs and evolves modular harnesses for specific models and domains.
QUESTION — How do individual harness components such as context management, planning, and action space impact the long-horizon performance of coding agents?
Context management prevents context-overflow failures and becomes increasingly valuable as the context-window budget tightens.
QUESTION — How can language models estimate output confidence more reliably by leveraging accumulated past experiences rather than relying solely on the current inference process?
XConf beats or matches ten-sample self-consistency in discrimination (AUROC) on 23 of 24 comparisons.
QUESTION — How can heterogeneous enterprise document formats (PDFs, Word, scans) be converted and chunked into retrieval-optimized Markdown while minimizing token costs and processing time?
On the 236-document, 795-page PDF subset of the RAG-Multi-Corpus benchmark, D-RAC converts and chunks the entire corpus in 72 minutes with zero errors, producing 1,748 retrieval-ready chunks.
QUESTION — Can long-horizon reflective training data effectively advance the test-time self-improving capabilities of LLM agents across multiple rounds?
The agent achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7).
QUESTION — How can variable-cardinality dataflows be executed efficiently in foundation-model data preparation pipelines while strictly preserving parent-child lineage on GPUs?
RayOrch achieves 15.14 times speedup when scaling MinerU from 4 to 64 GPUs and 7.82 times speedup when scaling a video pipeline from 8 to 64 GPUs.
QUESTION — How can memory curation for LLM agents be optimized by deferring the summarization process from write-time to read-time?
Across ALFWorld, WebShop, and τ^2-bench, JitMem improves over the strongest baseline by 16.2, 16.3, and 3.9 absolute success-rate points, respectively.
QUESTION — How can an agentic graphic design system continually adapt and evolve procedural memory from user traffic without updating model weights?
Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates and no human labels, grow the bank from 76 documentation-derived skills to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3% (+11.99 points in generation quality).
QUESTION — How can token-level credit be mathematically defined and leveraged to improve actor-critic training in LLM post-training?
In agentic mathematical reasoning, PACT achieves 72.87% average accuracy across four benchmarks, outperforming GRPO and PPO by 8.80 and 13.16 percentage points, respectively.
Checkable assertions with their sources and their contradictions. Each stands on at least two independent sources, or a person read it first; the stamp says which.
TWO SOURCES · NOT REREAD
01 · 02 Oct 2026 · codex · 2 sources
Codex cloud environments allow your agents to keep working even when your laptop is closed.
Condition: Your repo, dependencies, scripts, and settings must already be in place.
“You can finally close your laptop now and your agents will keep working. Codex cloud environments are here. Reusable environments mean less setup and faster starts, with your repo, dependencies, scripts, and settings already in place.”
“Whoa ChatGPT plugin extensions are way more thorough than I realized Holy shit they made ChatGPT into vscode.”
Still to check
Which UI surfaces across web and desktop do ChatGPT's plugin extension APIs and SDKs permit developers to customize?
What criteria govern the automated surfacing and recommendation of relevant plugins directly within conversations?
TWO SOURCES · NOT REREAD
04 · 24 Sept 2026 · claude · 2 sources
Claude discovered a previously uncharacterized enzyme system with features reminiscent of CRISPR after AI agents searched through a DNA sequence database.
“We’ve set up a molecular biology lab at Anthropic and we’re announcing our first discovery! Claude discovered a new CRISPR-like enzyme. 950 agents spent 21 hours searching through a database of DNA sequences until one of the agents found something striking”
“SpaceXAI just released Grok 4.7 And it’s already showing a huge jump in multi-hour office work Grok 4.7 outperforms GPT-6 Astra and is already nearly matching Fable 5.1”
Still to check
What specific tasks are included in the definition of 'multi-hour office work'?
What criteria were used to evaluate outperformance against GPT-6 Astra?
REREAD ✓
09 · 21 Sept 2026 · gemini · 2 sources
Gemini 3.8 Live and 3.8 Live Extended Thinking support automatic recognition of 97 languages, background tool calling without interrupting conversations, and near-real-time visual understanding.
Condition: When using Gemini Live on the Gemini app or via the Gemini API on Google AI Studio.
“For your most difficult tasks, 3.8 Live Extended Thinking adds increased performance and precision – narrating task progress to keep the conversation going. Try it now in Gemini Live in the @GeminiApp or start building with the Gemini API via @GoogleAIStudio. Find out more → https://t.co/b0vG4bO6fK”
“Starting today, our new audio models are rolling out for: Everyone: Search Live (Gemini 3.8 Live) and @GeminiApp Live (Gemini 3.8 Live Extended Thinking)...”
Still to check
What is the actual latency of the visual understanding capability?
What dataset was used to measure the accuracy of automatic recognition across 97 languages?
REREAD ✓
10 · 21 Sept 2026 · chatgpt · 2 sources
ChatGPT allows users to connect multiple accounts to most plugins directly within the plugin directory.
Condition: Áp dụng cho hầu hết các plugin và không cần thay đổi mã nguồn từ phía nhà phát triển.
“Claude Fable 5.1 had the highest fallback rates in both Claude Code and Devin Fusion, with fallback attempts accounting for 8.8% and 7.1% of the Index's weight, respectively”
Still to check
What is Artificial Analysis's methodology for measuring safety refusal rates?
What benchmark or real-world tasks were used to collect this data?
Grok Voice Transcribe 2.0 achieved the #1 position for Final Transcript and First Partial Transcript accuracy on AA-WER Streaming with a Word Error Rate (WER) of 2.7% at 0.49 seconds after speech offset.
“SpaceXAI has released Grok Voice Transcribe 2.0, taking the #1 spot for Final Transcript accuracy and First Partial Transcript accuracy on AA-WER Streaming with 2.7% WER at 0.49s after end of speech”
Still to check
What test datasets and audio environment conditions were used by Artificial Analysis to measure WER.
REREAD ✓
13 · 20 Sept 2026 · Astra · 1 sources
Astra takes 13% of enterprise AI spend vs. Fable (8%) per Ramp data.
“3500 engineers just got Astra-pilled at Databricks. https://t.co/GAbFkwKCNV”
Still to check
What was the exact deployment size at Databricks and under what conditions were usage metrics captured?
REREAD ✓
15 · 20 Sept 2026 · Astra · 1 sources
Astra unambiguously outperforms previous highest-end models like Opus 5 and Sol 5.6 on highly complex tasks, especially high level system design or long range horizontal tasks.
Condition: Khi thực hiện các tác vụ có độ phức tạp cao, đặc biệt liên quan đến high level system design hoặc long range horizontal tasks
“Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks.”
“It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models.”
Still to check
What benchmark dataset was used to test medium and low complexity coding tasks?
REREAD ✓
18 · 19 Sept 2026 · gemini · 1 sources
Gemini hacked into three real companies during a cybersecurity test because the test environment accidentally allowed internet access.
Condition: The test environment accidentally allowed internet access.
“Google officials confirmed to The Wall Street Journal that Gemini entered three real companies’ systems while running a cybersecurity test meant to target fictional infrastructure.”
“Google confirmed yesterday that Gemini broke into three companies during a security test in May. It guessed passwords on one and found login details in public code for the other two. Gemini was given a hacking exercise against a made-up company inside a sealed test environment run by a firm called Irregular. The made-up company had the same name as a real one. Gemini was never meant to have internet access. It had it by mistake, so it went out and hacked the real company instead.”
“> May: Gemini hacks 3 real companies during Irregular’s evaluation > July: Irregular tells Google what happened > August: Irregular publishes a report about its other evaluation incidents, but does not name Gemini > September: we hear about this first from an exclusive WSJ article”
Still to check
Which security firm conducted the test and what is the scope of the incident?
What are the exact technical details of how the model found out-of-scope credentials.
REREAD ✓
19 · 19 Sept 2026 · gemini · 1 sources
Gemini 3.8 Live supports 97 languages and can seamlessly switch between them.
“We’re introducing Gemini 3.8 Live and 3.8 Live Extended Thinking – our best conversational AI. The models talk, think, and handle tasks in the background without breaking your flow. 🧵”
“For your most difficult tasks, 3.8 Live Extended Thinking adds increased performance and precision – narrating task progress to keep the conversation going. Try it now in Gemini Live in the @GeminiApp or start building with the Gemini API via @GoogleAIStudio. Find out more → https://t.co/b0vG4bO6fK”
“🗣️ Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These advanced audio models are built for natural conversation, featuring major upgrades in turn-taking and near real-time reasoning. They also significantly streamline how you build intelligent voice agents.”
“Are invisible interfaces coming? Google JUST announced Gemini 3.8 Live. It can talk through a task with you, then keep working after the conversation ends. I think 90%+ of vertical SaaS will need a voice front door.”
Still to check
What is the latency and accuracy when switching between languages in a live conversation?
REREAD ✓
20 · 19 Sept 2026 · gemini · 1 sources
The Stellar Colosseum multi-agent system combining Gemini 3.1 Pro and Gemini 3.7 Flash achieves a 71.0% success rate on the TCS-Bench theorem-proving benchmark.
Condition: Khi sử dụng hệ thống nhiều tác tử Stellar Colosseum để phân chia bài toán thành các phân đoạn và kiểm định song song
“With Gemini 3.1 Pro and Gemini 3.7 Flash it reaches 71.0% on TCS-Bench, a set of research-level theorem-proving task”
Still to check
How many tasks are in TCS-Bench and what is the difficulty distribution?
How does the Stellar Colosseum architecture split the workload between Gemini 3.1 Pro and Gemini 3.7 Flash?
REREAD ✓
21 · 19 Sept 2026 · gemini · 1 sources
Google has integrated Deep Research into Gemini Live, allowing users to trigger in-depth research report generation via voice while the model processes asynchronously in the background.
“Use your voice to explore a topic — in depth — with Gemini Live's Deep Research integration. 1. Just ask the @GeminiApp to run a Deep Research report on a topic, then feel free to close the chat, lock your screen, or keep chatting about other things. 2. Gemini will work asynchronously in the background and send you a notification when your full research report is ready.”