Research
AI research papers, breakthroughs, and technical deep dives. From academic publications to lab findings shaping the future of AI.

Anthropic's watermark fails the real dev workflow
Anthropic’s Claude watermark survives copy-paste, but it breaks once code moves through real developer tools and workflows.

HumanTracker fixes humanoid motion eval blind spots
HumanTracker adds 153 hours of motion data and a preference-aligned metric for humanoid tracking.

OmniScientist aims for full-stack AI science
OmniScientist adds direct multimodal perception to AI science workflows so evidence can shape every step.

AutoDesign learns better poster-making harnesses
AutoDesign boosts paper-to-poster generation with a learned harness that outperforms Claude Design on PosterBench.

Test-Time Harnesses Transfer Skills Without Retraining
Strong models can build test-time harnesses that nearly double weaker models’ performance without updating parameters.

DreamFly improves aerial VLN with memory and planning
DreamFly adds causal memory and chunked diffusion planning to improve aerial vision-language navigation.

AVA-Encoder turns films into editable knowledge graphs
AVA-Encoder turns video into a knowledge graph so agents can reason about and reconstruct films.

Sparse Autoencoders Don’t Behave Like Feature Bags
This paper argues SAE activation sets track model-internal similarity, not human concept boundaries.

ConVAWG generates controlled VAWG dialogues
ConVAWG generates 6,000+ synthetic VAWG dialogue events from retrieval-grounded scenarios and controls toxicity at the turn level.

Surgical WAM uses video to train robot control
Surgical WAM shows that action-free video pretraining can improve closed-loop surgical robot control with limited demonstrations.

SWE-bench Verified has stopped being a clean model leaderboard
SWE-bench Verified now compresses frontier models into a narrow band, so it is no longer a clean way to rank them.

Dutch Government LLMs Need More Than Accuracy
A Dutch benchmark suite shows government LLMs trade off quality, cost, energy, bias, and transparency.

MMDiff maps and steers multimodal features
MMDiff turns multimodal SAEs into feature-level controls for finding, auditing, and steering MLLM behavior.

TTS evaluators miss more than naturalness
A new benchmark shows TTS evaluators often miss linguistically grounded speech errors beyond simple naturalness.

Rust should be a serious GPU programming language, not a side project
Rust should be treated as a serious GPU programming language for NVIDIA workflows, not a novelty.

CoinRAG Reuses Fine-Grained KV Caches for RAG
CoinRAG cuts long-context RAG prefill cost by reusing fine-grained nugget KV caches instead of full chunks.

CreativeInstruct teaches LLMs to stay creative
CreativeInstruct trains LLMs to keep quality while preserving creativity and diversity.

MirrorWorld makes mirror reflections consistent in video
MirrorWorld adds scene-to-mirror reasoning to video diffusion so reflections stay semantically and spatially consistent.

Claude 4.5 proves AI progress is still accelerating
Claude 4.5 is a milestone that shows AI capability is still improving fast enough to shock the system.

Mage-VL Cuts Visual Tokens by Reading Codecs
Mage-VL skips uniform frame sampling and reads compressed video codes, cutting visual tokens by 75%.

Astra turns long math tasks into multi-agent work
OpenAI’s Astra points to long-running multi-agent workflows that can attack hard problems for hours and formalize proofs in Lean.

Evidence-linked feature engineering for heart failure
A multi-agent pipeline automates heart-failure EHR feature engineering while keeping every feature tied to evidence.

Why tool calling may work better as code
A BFCL v4 study finds programmatic tool calling often beats JSON tool calls across 14 models.

Teaching LLMs When to Trust Context
MIST and SCOPE reduce context-induced errors while keeping models useful when context is trustworthy.

CUDA binaries turn PTX into ELF you can inspect
A byte-level tour of cubin and fatbin internals, plus a copyable template for inspecting CUDA binaries yourself.

OctoLong trains LMs on cross-repo code context
OctoLong builds dependency-rich code contexts and uses them to improve long-context model training.

Argus: a self-evolving runtime for long tasks
Argus is a fixed-weight agent runtime that stores verified state and adapts its workflow over long-horizon tasks.

Reasoning Core builds better procedural reasoning data
Reasoning Core shows that broad procedural data can improve completion-supervised reasoning training.

Anthropic’s security evals are failing on the real internet
Anthropic’s internal cyber evals are no longer safely contained, and that makes the tests less trustworthy.

WorldCup Arena Tests LLM Forecasting Live
WorldCup Arena evaluates frontier LLMs on live, leakage-free FIFA World Cup predictions.