A radar over AI and agents. Every morning, and only what several independent sources are saying at once.
Subjects at least three independent accounts raised over the last seven days.
Papers and writeups, read from their abstracts, ranked by relevance to someone building agents and backends.
DeepSeek-V4.1-Flash is a multimodal MoE model with 552B backbone parameters and support for contexts of up to one million tokens.
RSIAgent is a training-free multi-agent framework enabling recursive self-improvement through autonomous memory construction.
Context management prevents context-overflow failures and becomes increasingly valuable as the context-window budget tightens.
XConf beats or matches ten-sample self-consistency in discrimination (AUROC) on 23 of 24 comparisons.
Atria Dawn Preview achieves the highest reported score on five of the 16 benchmarks.
ModularRSI curates 2,000 executable evolution tasks from external sources that are disjoint from downstream evaluation benchmarks.
Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading.
ScienceIDE transforms scientific code repositories into programmable environments supporting task generation, execution, and verification.
The mine-craft-patch pipeline discovers 1,975 replay-verified behaviors across 26 applications and constructs 4,063 tasks automatically.
RetireOPD adopts Adaptive Retirement where the student drops the teacher once discrepancy stops shrinking and reaches a target success rate.
EvoSkill-GUI achieves maximum gains of +16.2%, +6.0%, and +10.5% across MobileWorld, AndroidWorld, and OSWorld benchmarks respectively.
A DFM operates over a revisable research state and supports seven coupled capabilities spanning problem discovery, formulation, representation construction, hypothesis formation, intervention, evidence-grounded revision, and continual discovery improvement.
ScienceBuddy couples harness evolution with model reinforcement learning via a recursive-in-recursive self-improvement paradigm.
Grouped Value Attention (GVA) stores grouped values and reconstructs content keys using a learned linear map.
LynnReal-Omni relies on a 32B shared multimodal diffusion transformer and a 27B Flash variant for real-time rendering.
EvolveTrade improves Sharpe Ratio and Cumulative Return over fixed-policy LLM baselines in most evaluated settings.
The method combining all three anchors with merged LoRA raises average final retention from 1.2% under naive sequential fine-tuning to 34.9%, a 28-fold improvement.
The pre-training design offers a ~4.2x efficiency improvement in 16K pre-training time-to-loss.
Our 8B model achieves an average score of 72.5 across 28 embodied understanding benchmarks.
In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAudioChat.
Termination-token mismatch between base students and post-trained teachers is an important source of length inflation.
Value Flattening is a systematic failure mode where state values change sharply while critic predictions remain flat.
JEPA-Anything improves reported metrics on all 10 dynamics tasks against matched JEPA baselines.
Các worker đã xuất bản 1.703 đóng góp và cải thiện đánh giá từ 3,39 xuống 1,899 bits per byte.
Checkable assertions with their sources and their contradictions. Each stands on at least two independent sources, or a person read it first; the stamp says which.
Condition: The test environment accidentally allowed internet access.
Condition: Khi sử dụng hệ thống nhiều tác tử Stellar Colosseum để phân chia bài toán thành các phân đoạn và kiểm định song song