A radar over AI and agents. Every morning, and only what several independent sources are saying at once.
Subjects at least three independent accounts raised over the last seven days.
Papers and writeups, read from their abstracts, ranked by relevance to someone building agents and backends.
AGP achieves success rates of 100%, 100%, and 80% on three block construction configurations.
SELF-INDEX enables an index to self-evolve without human intervention through autonomous diagnosis and key revision by its Optimizer.
On Qwen3-8B, it trades some pass@1 reliability for higher pass@k (+3.4 pp at pass@16) and solves more distinct tasks.
On A100 GPUs, eight serial calls use 4.64-4.86x as much gross GPU-device energy as one batched call with eight candidates.
HazardAuditor improves accuracy by up to 16.5 percentage points over the strongest prior guard.
VDN-H3 completes DiT denoising for a 14.3-second, 768p video in 6.70 seconds on eight NVIDIA B200 GPUs.
On AIME24, Pass@3 increases by 10.0% while token usage is reduced by 27.9% relative to the base model.
The PACT benchmark spans twelve regulated enterprise domains and forty-eight conversation scenarios.
It achieves up to 9.7% relative improvement in-domain.
The best-performing model scores 100.0% on target localization and 78.9% on spatial relations in the annotated run.
AMPLE-Math is a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer.
SpectralShift reparameterizes alpha projections initialization to reshape the decay spectrum for Gated DeltaNet.
Vision-RL2 improves accuracy over the base model at every token budget and surpasses its largest-budget accuracy with about 4 times fewer visual tokens.
FLAT achieves a T2I GenEval score of 71.1.
UFO achieves the highest correlation with human evaluation preferences, delivering an average improvement of 15.25%.
It lowers whole-object Chamfer distance by 40%.
VākQA provides 2,001 factoid question-answer pairs across six domains with 2.53 hours of speech audio in Telugu.
FAMOS predicts movable-part segmentation and joint parameters from a sparse, unordered set of partial point clouds.
Repositories climbing on GitHub today, scored on the day's momentum, rank, adoption, freshness and project health.
Checkable assertions with their sources and their contradictions. Each stands on at least two independent sources, or a person read it first; the stamp says which.
Condition: Khi thực hiện các tác vụ có độ phức tạp cao, đặc biệt liên quan đến high level system design hoặc long range horizontal tasks
Condition: suspecting those tasks are mostly saturated by existing models