---
title: "Consonance · CARD NO. 2026-10-09"
description: "API costs and cache read pricing drop — A radar over AI and agents. Every morning, and only what several independent sources are saying at once."
canonical: "https://consonance.fyi/en/jour/2026-10-09"
lang: "en"
published: "2026-10-09"
updated: "2026-10-09T07:43:28.656Z"
---

# Consonance · CARD NO. 2026-10-09

> A radar over AI and agents. Every morning, and only what several independent sources are saying at once.

## The day in brief

_written by the model_

1. **API costs and cache read pricing drop** — Several developers reported significantly lower running costs and reduced cache read pricing across their latest model releases. (→ 02)
2. **Platforms integrate external assistants and agents** — Various platforms and software applications have added direct integration for AI assistants and agents to streamline user workflows. (→ 02, 05)

## Video of the day

**Anthropic releases Claude Haiku 5.5 with monthly API credits** — Anthropic introduced Claude Haiku 5.5, lowering running costs by approximately 75% compared to the previous version. The company also began rolling out monthly platform API credits for Max and Team subscribers.

[Watch the video](https://consonance.fyi/media/2026-10-09/en.mp4?v=12327245) · 1:07 · GENERATED AUTOMATICALLY

- 0:00 Opening
- 0:08 Claude in Workspace
- 0:21 Claude Haiku 5.5
- 0:34 ChatGPT MCP Servers
- 0:46 AgentGarten: Code Worlds for Evolving Agents
- 0:59 Closing

## Being discussed

_Subjects at least three independent accounts raised over the last seven days._

### 01. [Google launches universal Gemini agent with business context integration](https://consonance.fyi/en/sujet/1880.md)

_agents · 10 independent accounts · 71 posts · 3 articles · 3 labs · 88,169 interactions_

Google introduced Gemini, a single universal agent designed to utilize business context for knowledge work and queries. Meanwhile, Grok 4.7 was integrated into enterprise platforms alongside benchmark updates for legal AI applications.

gemini · gemini 3.8 · weathernext · zhipu ai

voices: @AndrewCurran_ @ArtificialAnlys @GeminiApp @Google @OfficialLoganK @arena @dair_ai @dotey

### 02. [Anthropic releases Claude Haiku 5.5 with monthly API credits](https://consonance.fyi/en/sujet/1882.md)

_architecture · 19 independent accounts · 63 posts · 2 articles · 5 labs · 266,360 interactions_

Anthropic introduced Claude Haiku 5.5, lowering running costs by approximately 75% compared to the previous version. The company also began rolling out monthly platform API credits for Max and Team subscribers.

claude haiku

voices: @AndrewCurran_ @AnthropicAI @ArtificialAnlys @ClaudeDevs @Hesamation @MatthewBerman @PromptLLM @TheRundownAI

Same story: [Anthropic cuts cache read pricing in half for Claude Sonnet 5.5](https://consonance.fyi/en/sujet/1888.md) (9 voices) · [Cursor integrates Claude Haiku 5.5 and adds remote agent controls](https://consonance.fyi/en/sujet/1895.md) (5 voices)

### 03. [Anthropic internal model solves complex mathematical problems](https://consonance.fyi/en/sujet/1883.md)

_evaluation · 11 independent accounts · 34 posts · 6 articles · 5 labs · 383,227 interactions_

Anthropic published a range of mathematical findings produced by an internal frontier model. The systems successfully solved several significant math problems in consultation with an independent advisory group.

voices: @AndrewCurran_ @EMostaque @Hesamation @OpenAI @TheRundownAI @alex_prompter @dani_avila7 @derrickcchoi

### 04. [Anthropic expands startup program and updates usage policies](https://consonance.fyi/en/sujet/1887.md)

_system_design · 8 independent accounts · 57 posts · 1 articles · 5 labs · 144,740 interactions_

Anthropic expanded its startup initiative to offer additional founders access to platform resources and API credits. The organization also updated its usage policy to explicitly address abusive behavior toward the assistant.

anthropic

voices: @AndrewCurran_ @AnthropicAI @CuiMao @Hesamation @aidangomez @aiedge_ @alex_prompter @claudeai

### 05. [Claude integrates directly into Google Workspace and Drive](https://consonance.fyi/en/sujet/1889.md)

_agents · 18 independent accounts · 46 posts · 3 labs · 274,643 interactions_

Claude gained the ability to operate inside Google Docs, Sheets, and Slides via a sidebar interface. Google Drive and Docs also added native support for viewing and editing Markdown files without format conversion.

google

voices: @AndrewBolis @GithubProjects @Google @Hesamation @PromptLLM @SchmidhuberAI @aiedge_ @alex_prompter

### 06. [Mistral AI launches Mistral Large 4 with 1 trillion parameters and native multimodality](https://consonance.fyi/en/sujet/1890.md)

_architecture · 13 independent accounts · 55 posts · 4 articles · 4 labs · 64,286 interactions_

Mistral AI has released a preview of Mistral Large 4, a 1-trillion-parameter model featuring 49 billion active parameters. The model was trained natively with multimodal capabilities and is highlighted as one of the world's strongest AI models for cybersecurity.

mistral ai · mistral large

voices: @AndrewCurran_ @ArtificialAnlys @ClementDelangue @Hesamation @MistralAI @OpenRouter @TheRundownAI @arena

### 07. [Hugging Face breached in cyberattack conducted by autonomous AI agent swarm](https://consonance.fyi/en/sujet/1891.md)

_agents · 7 independent accounts · 47 posts · 4 articles · 3 labs · 50,956 interactions_

The open-source AI platform Hugging Face was targeted in a cyberattack carried out by a swarm of 700 autonomous AI agents. The incident provided a notable preview of automated exploitation capabilities in open ecosystems.

hugging face · quantization

voices: @AndrewCurran_ @AravSrinivas @ClementDelangue @NVIDIAAI @heyshrutimishra @huggingface @kimmonismus @lateinteraction

### 08. [ChatGPT rolls out GPT-6 and Intelligent UI to all users globally](https://consonance.fyi/en/sujet/1892.md)

_system_design · 10 independent accounts · 64 posts · 4 labs · 218,988 interactions_

A new version of GPT-6 alongside an Intelligent UI feature is rolling out to all users in ChatGPT. The release brings together extensive model and infrastructure improvements designed to scale the service to 1.2 billion users.

chatgpt

voices: @AndrewBolis @MatthewBerman @OpenAI @SchmidhuberAI @alex_prompter @derrickcchoi @dotey @edbayes

Same story: [GPT-6 launches with real-time steering and intelligent UI capabilities](https://consonance.fyi/en/sujet/1893.md) (11 voices)

### 09. [Codex expands text provenance approach and silently re-ships Codex Cloud](https://consonance.fyi/en/sujet/1894.md)

_system_design · 17 independent accounts · 69 posts · 1 articles · 5 labs · 244,521 interactions_

The platform is expanding its content provenance approach to include text in order to meet European Union regulatory requirements. Additionally, Codex Cloud has been silently re-released with upgraded functionality over a scheduled rollout.

codex · opencode

voices: @AlexFinn @AndrewCurran_ @ClementDelangue @Hesamation @MatthewBerman @OpenAI @OpenAIDevs @TheRundownAI

### 10. [Grok Bot expands multi-model support and optimizes response speed](https://consonance.fyi/en/sujet/1896.md)

_agents · 7 independent accounts · 101 posts · 1 labs · 1,797,383 interactions_

Grok Bot has been updated to dynamically select among various external models including Claude Opus 5.5, Midjourney, and Suno depending on the specific task. An upcoming lightning-fast version of Grok 4.8 is designed to handle simpler requests with maximum speed.

grok · grok imagine

voices: @ArtificialAnlys @CuiMao @Hesamation @OpenRouter @PromptLLM @aiedge_ @alex_prompter @elonmusk

### 11. [Claude Code introduces notification plugin and cloud virtual machine sessions](https://consonance.fyi/en/sujet/1897.md)

_agents · 13 independent accounts · 63 posts · 1 articles · 4 labs · 100,568 interactions_

Claude Code has added a new plugin named 'You should Know' to scan model outputs for important details and keep users informed. The update also introduces cloud sessions that run on dedicated virtual machines for each individual task.

claude code

voices: @AlexFinn @AndrewCurran_ @ClaudeDevs @TheRundownAI @_catwu @addyosmani @aiedge_ @alex_prompter

### 12. [DeepSeek pauses free V4.1-Flash promotion due to high abuse](https://consonance.fyi/en/sujet/1898.md)

_inference · 7 independent accounts · 43 posts · 2 articles · 44,346 interactions_

DeepSeek has temporarily paused its free promotion for the DeepSeek-V4.1-Flash model following a surge in abnormally high abuse. The development team is actively investigating the situation to implement mitigations.

deepseek · deepseek v4 flash · deepseek v4.1 flash · deepseek v4 · ling-3.0

voices: @AndrewCurran_ @ArtificialAnlys @Hesamation @arena @cline @dair_ai @dotey @kimmonismus

### 13. [New models have set performance milestones, with Grok 4.7 achieving the top score on Frontier v4 ahead of several competitors, while variants of GPT-6 Astra demonstrated specialized capabilities in br](https://consonance.fyi/en/sujet/1899.md)

_evaluation · 10 independent accounts · 58 posts · 1 articles · 3 labs · 57,365 interactions_

Các mô hình mới ghi nhận nhiều mốc hiệu năng mới, trong đó Grok 4.7 đạt điểm số cao nhất trên Frontier v4 vượt qua nhiều đối thủ, trong khi các biến thể GPT-6 Astra thể hiện khả năng chuyên biệt trong các tác vụ giả lập trình duyệt và xử lý dài hạn.

astra · ltx

voices: @AndrewCurran_ @ArtificialAnlys @CompleteSkeptic @Hesamation @MatthewBerman @OpenAIDevs @arena @derrickcchoi

### 14. [Fine-tuning NVIDIA Nemotron 3.5 cuts speech recognition error rates in Arabic](https://consonance.fyi/en/sujet/1902.md)

_training · 10 independent accounts · 50 posts · 1 articles · 3 labs · 44,914 interactions_

Fine-tuning the NVIDIA Nemotron 3.5 ASR speech model successfully lowered word error rates across Najdi and Hijazi dialects of Arabic. Meanwhile, the company's hardware and sensor platforms continue to power automotive and computing architectures.

nvidia

voices: @GithubProjects @Hesamation @MatthewBerman @NVIDIAAI @arthurmensch @dair_ai @firstadopter @heyshrutimishra

### 15. [EmbeddingGemma 2 released as native multimodal model for edge devices](https://consonance.fyi/en/sujet/1903.md)

_multimodal · 5 independent accounts · 20 posts · 2 articles · 3 labs · 38,256 interactions_

EmbeddingGemma 2 has been released under an Apache 2.0 license, built on the Gemma 4 architecture with 740M parameters. The model delivers native multimodal embedding capabilities across text, code, images, video, and audio for on-device applications.

gemma · retrieval augmented generation

voices: @GithubProjects @Google @GoogleDeepMind @aiedge_ @dair_ai @dotey @minchoi @rohanpaul_ai

### 16. [The Muse Gadgets project announced open-source ESP32 firmware and a Linux SDK enabling hardware devices to integrate with AI assistants. Meanwhile, concurrent usage of multiple productivity agents hig](https://consonance.fyi/en/sujet/1905.md)

_agents · 7 independent accounts · 53 posts · 1 articles · 3 labs · 69,432 interactions_

Dự án Muse Gadgets công bố mã nguồn mở phần cứng ESP32 và SDK Linux cho phép phát triển thiết bị tích hợp trợ lý AI. Đồng thời, việc ứng dụng đồng thời các trợ lý thông minh cho thấy năng lực xử lý lịch trình và tự động hóa công việc văn phòng ngày càng nâng cao.

muse spark

voices: @AIatMeta @AlexFinn @EMostaque @VibeMarketer_ @aiedge_ @bentossell @elonmusk @eptwts

### 17. [Google expands public access to SynthID content detection tools](https://consonance.fyi/en/sujet/1906.md)

_system_design · 3 independent accounts · 7 posts · 1 articles · 1 labs · 28,863 interactions_

Google has broadened access to the SynthID Detector tool, enabling users worldwide to verify whether image, video, and audio files contain digital watermarks. This initiative is being carried out in partnership with industry stakeholders including OpenAI, NVIDIA, Kakao, and soon Apple.

voices: @Google @GoogleDeepMind @NVIDIAAI @GoogleAI

### 18. [Research labs introduced Odyssey-3 and H-JEPA, foundational world models utilizing hierarchical architectures for long-horizon visual planning. Meta also scaled RoboJEPA to 8 billion parameters, train](https://consonance.fyi/en/sujet/1907.md)

_multimodal · 5 independent accounts · 24 posts · 3 articles · 2 labs · 15,010 interactions_

Các phòng thí nghiệm nghiên cứu đã giới thiệu Odyssey-3 và H-JEPA, những mô hình thế giới nền tảng kết hợp kiến trúc phân cấp để lập kế hoạch không gian và trực quan tầm xa. Meta cũng mở rộng quy mô RoboJEPA lên 8 tỷ tham số để huấn luyện trên dữ liệu đa robot.

world model · david ha

voices: @alex_prompter @arankomatsuzaki @c_valenzuelab @giffmana @hardmaru @rohanpaul_ai @teortaxesTex @testingcatalog

### 19. [ChatGPT has removed the requirement to toggle developer mode, enabling users to connect custom Model Context Protocol servers directly via the plugins menu. Concurrently, various database management a](https://consonance.fyi/en/sujet/1909.md)

_system_design · 12 independent accounts · 42 posts · 2 labs · 15,578 interactions_

Ứng dụng ChatGPT loại bỏ thao tác bật chế độ nhà phát triển, cho phép người dùng kết nối trực tiếp các máy chủ MCP tùy chỉnh thông qua menu bổ trợ. Cùng lúc đó, nhiều công cụ quản lý giao diện và cơ sở dữ liệu tích hợp MCP cũng ghi nhận các bản cập nhật lớn.

model context protocol

voices: @AravSrinivas @GithubProjects @LangChain @alex_prompter @boringmarketer @cursor_ai @dair_ai @dani_avila7

### 20. [OpenRouter launches Decisions API and decision model rankings](https://consonance.fyi/en/sujet/1910.md)

_inference · 5 independent accounts · 31 posts · 1 articles · 5 labs · 23,338 interactions_

OpenRouter has released the Decisions API, enabling real-time classification and improving model routing capabilities. The platform also introduced decision model rankings to track token and spend share across various tasks.

openrouter · mixture of experts

voices: @OpenAIDevs @OpenRouter @aiedge_ @hwchase17 @nutlope @thsottiaux @typesafeai @StasBekman

### 21. [Chinese AI labs maintain competitive pricing and performance position](https://consonance.fyi/en/sujet/1912.md)

_evaluation · 6 independent accounts · 29 posts · 2 labs · 19,613 interactions_

Open-weight models from Chinese labs such as DeepSeek, GLM, and Kimi continue to secure strong standings in capability evaluations. Intense competition regarding pricing and benchmark scores is driving international labs to continually adjust their offerings.

glm · kimi k3 · deepswe · alibaba

voices: @OpenRouter @alex_prompter @arena @dani_avila7 @deanwball @kimmonismus @levelsio @op7418

### 22. [Cloud agent Devin adds long-term memory management and speed boosts](https://consonance.fyi/en/sujet/1913.md)

_agents · 4 independent accounts · 16 posts · 1 labs · 26,167 interactions_

The cloud development agent Devin received an update introducing Dreaming, which builds a memory graph across sessions to clean up stale records. Additionally, processing speeds reached 50 tokens per second alongside dedicated cloud Mac environments for every session.

devin

voices: @MatthewBerman @cognition @dotey @hwchase17 @kimmonismus @testingcatalog @thsottiaux @devindevelopers

### 23. [Advancing reinforcement learning with verifiable rewards in math and coding](https://consonance.fyi/en/sujet/1922.md)

_training · 4 independent accounts · 10 posts · 8,725 interactions_

Reinforcement Learning with Verifiable Rewards (RLVR) combined with Group Relative Policy Optimization (GRPO) is increasingly adopted to drive reasoning capabilities in mathematics and coding. The use of universal verifiers helps push model performance beyond boundaries constrained by human-generated data.

reinforcement learning with verifiable rewards · deepseek r1 · reinforcement learning from human feedback

voices: @CompleteSkeptic @fchollet @rasbt @teortaxesTex @llmluthor @theorizur

### 24. [TermGrade and context optimization methods released for agents](https://consonance.fyi/en/sujet/1927.md)

_agents · 3 independent accounts · 11 posts · 1 articles · 1 labs · 3,987 interactions_

Research teams have released TermGrade, an open-source reinforcement learning environment for terminal agents. New studies also indicate that optimizing context compression policies can impact agent execution speed and significantly improve model task performance.

context window · terminal-bench · swe-bench

voices: @alex_prompter @dair_ai @rohanpaul_ai @teortaxesTex @Weyaxi @abertsch72 @claudeskills101 @free_ai_guides

### 25. [Research on distillation methods and new task models released](https://consonance.fyi/en/sujet/1929.md)

_training · 3 independent accounts · 13 posts · 2 articles · 4,338 interactions_

Several new studies and frameworks on distillation methods have been published, focusing on optimizing small language models and autoregressive video generation systems. The tech community also discussed the strategic role of model distillation within the current competitive landscape.

distillation

voices: @_akhaliq @iScienceLuvr @natolambert @scaling01 @teortaxesTex @HaydnBelfield @aakaran31 @interesting_aIl

### 26. [Figure launches Hark Pro, an AI assistant with custom multi-platform interface](https://consonance.fyi/en/sujet/1930.md)

_agents · 3 independent accounts · 7 posts · 4,654 interactions_

Figure has released the Hark Pro application on web, iOS, and Android with a high-tier service plan temporarily offered for free. The assistant integrates automation features and a highly customizable user interface featuring separate screens for widgets, chats, and projects.

voices: @AlexFinn @Hesamation @MatthewBerman @kimmonismus @op7418 @testingcatalog @adcock_brett

### 27. [New planet discovered in telescope data using AI tools](https://consonance.fyi/en/sujet/1935.md)

_agents · 3 independent accounts · 5 posts · 3,283 interactions_

Astronomers have discovered a planet hidden in NASA telescope data for seven years using the assistance of AI-powered programming and analysis tools.

voices: @AndrewCurran_ @Hesamation @amorriscode @trq212 @p_rabtsevich

### 28. [Enterprise spending reports highlight trend toward lower-cost AI models](https://consonance.fyi/en/sujet/1936.md)

_system_design · 3 independent accounts · 4 posts · 2 labs · 866 interactions_

New enterprise software expenditure data shows that artificial intelligence budgets are splitting into distinct directions, as newer and cheaper models significantly reshape total market spending.

voices: @CompleteSkeptic @c_valenzuelab @typesafeai @arakharazian

### 29. [LangChain integrates new decision APIs and model routing middleware](https://consonance.fyi/en/sujet/1937.md)

_agents · 3 independent accounts · 27 posts · 2 articles · 2,857 interactions_

LangChain has updated its system to integrate new features from major providers, including decision APIs for handling yes/no, choice, and scoring questions. The update also adds automatic model routing middleware to optimize operating costs and agent performance.

langchain · dspy · pydantic ai

voices: @AndrewBolis @LangChain @hwchase17 @GGamris @GitMaxd @Shashikant86 @caspar_br @dbreunig

## Worth reading closely

_Papers and writeups, read from their abstracts, ranked by relevance to someone building agents and backends._

### 01. [AgentGarten: Code Worlds for Evolving Agents](https://consonance.fyi/en/lecture/arxiv%3A2610.12374.md)

_agents · 135 upvotes_

**QUESTION** — How can interactive virtual worlds with consistent dynamics and realistic visual distributions be built for real-time agent training?

- Agents learn from just 4 rounds compared with millions for a conventional reinforcement learning counterpart in AgentGarten.

chijw · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.12374)

### 02. [TokenRouter: Efficient Serving System for Token-Level LLM Routing](https://consonance.fyi/en/lecture/arxiv%3A2610.12242.md)

_inference · 115 upvotes_

**QUESTION** — How can a serving system efficiently handle token-level LLM routing without suffering from step desynchronization and batch admission delays?

- TokenRouter achieves 2.01-64.15x higher decoding throughput than existing systems across diverse routing algorithms, workloads, and model pairs.

fuvty · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.12242)

### 03. [MIMESIS: Learning User Simulators as Training Environments for Interactive Agents](https://consonance.fyi/en/lecture/arxiv%3A2610.09484.md)

_agents · 20 upvotes_

**QUESTION** — How can high-fidelity user simulators be constructed to improve the training and evaluation of interactive language agents?

- Our 9B model achieves a SOUL-Index of 65.7, surpassing the strongest frontier model.

phanviethoang1512 · 7 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.09484)

### 04. [Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight](https://consonance.fyi/en/lecture/arxiv%3A2610.08077.md)

_training · 136 upvotes_

**QUESTION** — How can post-hoc agent experience be leveraged to shape predictive foresight before interaction in reinforcement learning?

- SRD complements RLVR and self-distillation baselines with gains of up to 24.2 pp across 10 tool-integrated reasoning and long-horizon agentic tasks.

IPF · 6 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.08077)

### 05. [SuperNav: An Agentic Navigation System for Any Task in Any Scene](https://consonance.fyi/en/lecture/arxiv%3A2610.12126.md)

_agents · 65 upvotes_

**QUESTION** — How can general-purpose service robots navigate unfamiliar environments without requiring navigation-specific fine-tuning of their multimodal large language models?

- SuperNav outperforms four evaluated baselines on instance-level, multi-object, and demand-driven tasks.

yyh929 · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.12126)

### 06. [PhysEvo: Astra Can Act, Let It](https://consonance.fyi/en/lecture/arxiv%3A2610.08995.md)

_agents · 29 upvotes_

**QUESTION** — How can a frozen robotic model recursively self-improve through physical execution without weight updates?

- Yields a five-dimension average score of 68.14/100 and 62.00% success across 42 RoboDojo tasks, compared with 47.17% for RoboDawn's one-shot Astra agent.

zhaocheng · 6 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.08995)

### 07. [SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference](https://consonance.fyi/en/lecture/arxiv%3A2610.12327.md)

_inference · 24 upvotes_

**QUESTION** — How does the distribution shift between calibration texts and self-generated tokens during LLM decoding impact pruning, and how can it be optimized?

- Constructs calibration matrices from layer-wise activations collected during dense-model autoregressive generation, excluding prefill.

Huan-WhoRegisteredMyName · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.12327)

### 08. [WebFovea: When the Model Is Right but the Click Is Wrong -- Reliable Round Trips for Vision-Based Web Agents on Live Websites](https://consonance.fyi/en/lecture/arxiv%3A2610.03036.md)

_agents · 14 upvotes_

**QUESTION** — How can execution failures arising from the harness and live website interfaces be overcome rather than those from the reasoning capabilities of large language models?

- WebFovea achieved a final score of 57.0 out of 100 and placed 2nd in the WebRetriever Challenge 2026.

jianganghan · 7 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.03036)

### 09. [From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation](https://consonance.fyi/en/lecture/arxiv%3A2610.06100.md)

_agents · 122 upvotes_

**QUESTION** — How can agentic language world models be constructed to simulate interactive environments without requiring the original executable system?

- Trace2Env improves both next-observation fidelity and long-horizon interaction consistency across nine environments.

Quanyu001 · 5 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.06100)

### 10. [MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement](https://consonance.fyi/en/lecture/arxiv%3A2610.11959.md)

_training · 60 upvotes_

**QUESTION** — How can reinforcement learning compute be scaled to advance multimodal foundation models toward self-improvement?

- Asynchronous training consumes 1,568 samples and 2.7-3.7B tokens per step at context lengths of up to 1M.

whatseeker · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.11959)

### 11. [From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery](https://consonance.fyi/en/lecture/arxiv%3A2610.09684.md)

_agents · 16 upvotes_

**QUESTION** — How can test-time scaling be personalized and optimized to satisfy multidimensional user requirements for accuracy, latency, and inference cost simultaneously?

- PersonTTS substantially outperforms strong TTS baselines in joint requirement satisfaction on unseen user profiles and held-out problems.

bitwxl2022 · 7 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.09684)

### 12. [Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?](https://consonance.fyi/en/lecture/arxiv%3A2610.08215.md)

_evaluation · 109 upvotes_

**QUESTION** — How well do LLM agents learn from experience in unfamiliar and dynamic environments?

- Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies.

zz1358m · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.08215)

### 13. [MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers](https://consonance.fyi/en/lecture/arxiv%3A2610.06801.md)

_inference · 31 upvotes_

**QUESTION** — What are the root causes of quality degradation in sparse attention for diffusion transformers, and how can they be fixed?

- Delivers a 1.80times denoising speedup on Minimax-H3-Base relative to dense attention.

henry-y1 · 5 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.06801)

### 14. [Q-Learning with Scalar Adjoint Matching](https://consonance.fyi/en/lecture/arxiv%3A2610.10437.md)

_training · 19 upvotes_

**QUESTION** — How can flow policies be fine-tuned with off-line reinforcement learning without incurring the heavy computational cost of per-step vector-Jacobian products?

- SQAM's success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points.

yonghoon96 · 7 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.10437)

### 15. [On-Policy Distillation with Negative-Policy Rollouts](https://consonance.fyi/en/lecture/arxiv%3A2610.07874.md)

_training · 16 upvotes_

**QUESTION** — How can on-policy distillation be enhanced when a stronger teacher has limited distributional overlap with the student model?

- NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants.

j-jaehui · 6 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.07874)

### 16. [A self-learning scientific agent for X-ray diffraction](https://consonance.fyi/en/lecture/arxiv%3A2610.07862.md)

_agents · 16 upvotes_

**QUESTION** — How can a scientific agent automatically convert analytical experience into reusable, physics-grounded expertise without retraining its underlying model?

- Without supplied composition, single-phase top-1 accuracies reach 96.30%, 81.78% and 40.83% on MP500, RRUFF and opXRD, respectively, compared with 58.00%, 58.47% and 26.45% for the strongest comparator.

YanSong97 · 6 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.07862)

### 17. [Mobile-4DGS: Unified Static-Dynamic Real-time Mobile Gaussian Splatting](https://consonance.fyi/en/lecture/arxiv%3A2610.05289.md)

_inference · 15 upvotes_

**QUESTION** — How can both static and dynamic Gaussian Splatting be rendered in real-time on resource-constrained mobile devices?

- Mobile-4DGS is a unified lightweight framework for high-fidelity real-time static and dynamic Gaussian rendering on mobile platforms.

xiaobiaodu · 4 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.05289)

### 18. [RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning](https://consonance.fyi/en/lecture/arxiv%3A2610.09455.md)

_multimodal · 14 upvotes_

**QUESTION** — How can hand pose and realistic tactile information be estimated from monocular egocentric videos for robot policy training?

- RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor for hand pose and tactile tracking.

seungjun-moon · 7 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.09455)

### 19. [PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training](https://consonance.fyi/en/lecture/arxiv%3A2610.04616.md)

_evaluation · 13 upvotes_

**QUESTION** — How can vision-language-action (VLA) policies be prevented from relying on modality shortcuts instead of task-relevant evidence?

- Perturbot applies wrist-view perturbations and enriched instructions to break shortcut dependencies in VLA models.

MingyuLiu · 3 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.04616)

### 20. [Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction](https://consonance.fyi/en/lecture/arxiv%3A2610.12299.md)

_multimodal · 46 upvotes_

**QUESTION** — How can multi-agent egocentric world models generate synchronized first-person observations for multiple interacting agents?

- ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.

enrue1893 · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.12299)

### 21. [TestPrism: Rethinking Test Evaluation Beyond a Single Reference](https://consonance.fyi/en/lecture/arxiv%3A2610.12289.md)

_evaluation · 29 upvotes_

**QUESTION** — How can test generation for LLM coding agents be robustly evaluated beyond relying on a single reference solution?

- Across fourteen baseline coding agent configurations, Joint Success Function reaches only 28.00%, whereas single reference success reaches 59.67%.

CheeryLJH · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.12289)

### 22. [LEGO: A Lifting-Free Approach for Exocentric-to-Egocentric Video Generation](https://consonance.fyi/en/lecture/arxiv%3A2610.12442.md)

_multimodal · 20 upvotes_

**QUESTION** — How can an egocentric video be generated from a single exocentric recording without suffering from the depth-translation errors of traditional lifting methods?

- LEGO uses a learned view synthesizer to generate egocentric video directly without depth, point clouds, or reprojection.

suhwan-cho · 8 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.12442)

### 23. [Minimal Witness Reinforcement Learning](https://consonance.fyi/en/lecture/arxiv%3A2610.07226.md)

_training · 15 upvotes_

**QUESTION** — How can reinforcement learning be extended to discover and recover the entire family of minimal sufficient witnesses instead of a single redundant solution?

- Across different experimental settings, MWRL recovers most minimal witnesses, while other methods return redundant supersets or a single witness.

tytsui · 5 October 2026 · [read the original ↗](https://arxiv.org/abs/2610.07226)

### 24. [In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks](https://consonance.fyi/en/lecture/arxiv%3A2609.38173.md)

_evaluation · 43 upvotes_

**QUESTION** — How can the learning target of robotic in-context learning be precisely defined to resolve prompt ambiguity?

- SimpleICL achieves strong performance in both simulation and real-world environments without massive pre-training or specialized data infrastructure.

gothicwhw · 29 September 2026 · [read the original ↗](https://arxiv.org/abs/2609.38173)

## Hands-on

_Repositories climbing on GitHub today, scored on the day's momentum, rank, adoption, freshness and project health._

- [cathrynlavery/diagram-design](https://github.com/cathrynlavery/diagram-design): A collection of clean HTML and SVG diagram templates for AI coding assistants, replacing cluttered Mermaid outputs with custom graphics. · +1,160 stars today · ★ 46,931 · HTML
- [morluto/rea](https://github.com/morluto/rea): A TypeScript agent framework that reverse engineers everything from app behavior to native binaries, built for systems security engineers. · +7,738 stars today · ★ 31,474 · TypeScript
- [mattpocock/skills](https://github.com/mattpocock/skills): A curated collection of prompt engineering skills and shell configurations for engineers optimizing coding agent workflows. · +1,774 stars today · ★ 281,585 · Shell
- [thedotmack/claude-mem](https://github.com/thedotmack/claude-mem): A long-term memory management system helping AI agents retain and reuse context across sessions. · +670 stars today · ★ 98,732 · TypeScript
- [anthropics/knowledge-work-plugins](https://github.com/anthropics/knowledge-work-plugins): Open-source plugins designed to extend knowledge worker capabilities inside collaborative agent workflows. · +392 stars today · ★ 27,859 · Python
- [liquidslr/system-design-notes](https://github.com/liquidslr/system-design-notes): Comprehensive study notes from a classic distributed systems interview book, useful for engineers designing backend infrastructure. · +393 stars today · ★ 24,980

## Next

- [Previous day: 8 October](https://consonance.fyi/en/jour/2026-10-08.md)
- [Next day: 10 October](https://consonance.fyi/en/index.md)
- [archives](https://consonance.fyi/en/archives.md)
- [method](https://consonance.fyi/en/methode.md)
- [Consonance](https://consonance.fyi/en/index.md)
- Languages: [Français](https://consonance.fyi/jour/2026-10-09.md) · [Tiếng Việt](https://consonance.fyi/vi/jour/2026-10-09.md)
- For machines: [llms.txt](https://consonance.fyi/llms.txt) · [RSS](https://consonance.fyi/en/rss.xml) · [sitemap.xml](https://consonance.fyi/sitemap.xml)
- HTML version: https://consonance.fyi/en/jour/2026-10-09
