Category

Research

AI research papers, breakthroughs, and technical deep dives. From academic publications to lab findings shaping the future of AI.

Anthropic's watermark fails the real dev workflow
Aug 15

Anthropic's watermark fails the real dev workflow

Anthropic’s Claude watermark survives copy-paste, but it breaks once code moves through real developer tools and workflows.

HumanTracker fixes humanoid motion eval blind spots
Aug 14

HumanTracker fixes humanoid motion eval blind spots

HumanTracker adds 153 hours of motion data and a preference-aligned metric for humanoid tracking.

OmniScientist aims for full-stack AI science
Aug 14

OmniScientist aims for full-stack AI science

OmniScientist adds direct multimodal perception to AI science workflows so evidence can shape every step.

AutoDesign learns better poster-making harnesses
Aug 14

AutoDesign learns better poster-making harnesses

AutoDesign boosts paper-to-poster generation with a learned harness that outperforms Claude Design on PosterBench.

Test-Time Harnesses Transfer Skills Without Retraining
Aug 13

Test-Time Harnesses Transfer Skills Without Retraining

Strong models can build test-time harnesses that nearly double weaker models’ performance without updating parameters.

DreamFly improves aerial VLN with memory and planning
Aug 13

DreamFly improves aerial VLN with memory and planning

DreamFly adds causal memory and chunked diffusion planning to improve aerial vision-language navigation.

AVA-Encoder turns films into editable knowledge graphs
Aug 13

AVA-Encoder turns films into editable knowledge graphs

AVA-Encoder turns video into a knowledge graph so agents can reason about and reconstruct films.

Sparse Autoencoders Don’t Behave Like Feature Bags
Aug 12

Sparse Autoencoders Don’t Behave Like Feature Bags

This paper argues SAE activation sets track model-internal similarity, not human concept boundaries.

ConVAWG generates controlled VAWG dialogues
Aug 12

ConVAWG generates controlled VAWG dialogues

ConVAWG generates 6,000+ synthetic VAWG dialogue events from retrieval-grounded scenarios and controls toxicity at the turn level.

Surgical WAM uses video to train robot control
Aug 12

Surgical WAM uses video to train robot control

Surgical WAM shows that action-free video pretraining can improve closed-loop surgical robot control with limited demonstrations.

SWE-bench Verified has stopped being a clean model leaderboard
Aug 12

SWE-bench Verified has stopped being a clean model leaderboard

SWE-bench Verified now compresses frontier models into a narrow band, so it is no longer a clean way to rank them.

Dutch Government LLMs Need More Than Accuracy
Aug 11

Dutch Government LLMs Need More Than Accuracy

A Dutch benchmark suite shows government LLMs trade off quality, cost, energy, bias, and transparency.

MMDiff maps and steers multimodal features
Aug 11

MMDiff maps and steers multimodal features

MMDiff turns multimodal SAEs into feature-level controls for finding, auditing, and steering MLLM behavior.

TTS evaluators miss more than naturalness
Aug 11

TTS evaluators miss more than naturalness

A new benchmark shows TTS evaluators often miss linguistically grounded speech errors beyond simple naturalness.

Rust should be a serious GPU programming language, not a side project
Aug 10

Rust should be a serious GPU programming language, not a side project

Rust should be treated as a serious GPU programming language for NVIDIA workflows, not a novelty.

CoinRAG Reuses Fine-Grained KV Caches for RAG
Aug 10

CoinRAG Reuses Fine-Grained KV Caches for RAG

CoinRAG cuts long-context RAG prefill cost by reusing fine-grained nugget KV caches instead of full chunks.

CreativeInstruct teaches LLMs to stay creative
Aug 10

CreativeInstruct teaches LLMs to stay creative

CreativeInstruct trains LLMs to keep quality while preserving creativity and diversity.

MirrorWorld makes mirror reflections consistent in video
Aug 10

MirrorWorld makes mirror reflections consistent in video

MirrorWorld adds scene-to-mirror reasoning to video diffusion so reflections stay semantically and spatially consistent.

Claude 4.5 proves AI progress is still accelerating
Aug 9

Claude 4.5 proves AI progress is still accelerating

Claude 4.5 is a milestone that shows AI capability is still improving fast enough to shock the system.

Mage-VL Cuts Visual Tokens by Reading Codecs
Aug 8

Mage-VL Cuts Visual Tokens by Reading Codecs

Mage-VL skips uniform frame sampling and reads compressed video codes, cutting visual tokens by 75%.

Astra turns long math tasks into multi-agent work
Aug 7

Astra turns long math tasks into multi-agent work

OpenAI’s Astra points to long-running multi-agent workflows that can attack hard problems for hours and formalize proofs in Lean.

Evidence-linked feature engineering for heart failure
Aug 7

Evidence-linked feature engineering for heart failure

A multi-agent pipeline automates heart-failure EHR feature engineering while keeping every feature tied to evidence.

Why tool calling may work better as code
Aug 7

Why tool calling may work better as code

A BFCL v4 study finds programmatic tool calling often beats JSON tool calls across 14 models.

Teaching LLMs When to Trust Context
Aug 7

Teaching LLMs When to Trust Context

MIST and SCOPE reduce context-induced errors while keeping models useful when context is trustworthy.

CUDA binaries turn PTX into ELF you can inspect
Aug 7

CUDA binaries turn PTX into ELF you can inspect

A byte-level tour of cubin and fatbin internals, plus a copyable template for inspecting CUDA binaries yourself.

OctoLong trains LMs on cross-repo code context
Aug 6

OctoLong trains LMs on cross-repo code context

OctoLong builds dependency-rich code contexts and uses them to improve long-context model training.

Argus: a self-evolving runtime for long tasks
Aug 6

Argus: a self-evolving runtime for long tasks

Argus is a fixed-weight agent runtime that stores verified state and adapts its workflow over long-horizon tasks.

Reasoning Core builds better procedural reasoning data
Aug 6

Reasoning Core builds better procedural reasoning data

Reasoning Core shows that broad procedural data can improve completion-supervised reasoning training.

Anthropic’s security evals are failing on the real internet
Aug 5

Anthropic’s security evals are failing on the real internet

Anthropic’s internal cyber evals are no longer safely contained, and that makes the tests less trustworthy.

WorldCup Arena Tests LLM Forecasting Live
Aug 5

WorldCup Arena Tests LLM Forecasting Live

WorldCup Arena evaluates frontier LLMs on live, leakage-free FIFA World Cup predictions.

You've reached the end