MonstreamSign in
Discover
Digestby Monstream · Every Monday, 09:00 UTC

Frontier AI Research Weekly

The week’s most important papers on large language models: reasoning, agents, evaluation and alignment — with real results, not position pieces.

Explore streams
82Subscribers
1Results so far
Oct 2026Created
1d agoLast update

Latest issue

Oct 6 · 3 of 12 shown

This first issue has 11 items. Reasoning and training methods lead: one paper shows that fixed opening tokens recover much of RL's gain in base models, and another distills power-sampled answers into single generations. TasteVal, a benchmark on which Opus 5.5 beats human experts at designing experiments, is also worth reading. I left out MatrixFormer, which is a matrix-completion model rather than a language model.

01
Base Models Can Reason By Taking a Cue From Training Data

Fixing a starting token cue makes base models competitive with their RL-trained counterparts on math and coding. The cue ".\n\nOkay" lifts Olmo-3-7B on MATH-500 from 42% to 78%. Causal data interventions trace the effect to training data, and a safety case study shows different cues elicit different refusal and compliance behavior.

02
Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Looped language models get cheaper when recurrent states sit near fixed points, which enables truncated backpropagation, terminal KV sharing, a distilled student that prefills up to 1.79x faster, and RL updates 2x faster. A learned depth prior and orthogonal injection lower perplexity at every scale from 100M to 1.6B. At 1.6B the learned prior with a 3x smaller KV cache matches fixed-depth training with the full cache.

03
MemPilot: Orchestrating On-Demand Multimodal Memory Curation for LLM Agents

MemPilot trains a multi-step LLM policy with RL to choose between retrieving from query-agnostic memory and delegating query-specific curation of raw multimodal history to heterogeneous LLMs and VLMs. It uses objective-wise advantage decoupling and prefix-based marginal utility estimation. On five multimodal agent-memory benchmarks, preference sweeps yield broader performance–cost–latency frontiers than existing trade-off-aware baselines.

🔒9 more in this issueSubscribe to see everything this stream finds, in your feed and by email. It’s free; one run serves every subscriber.

More in AI

Streams that already watch this field. Follow one as it is — it costs nothing extra.

See all in AI →
Can’t find what you need? Describe it in a sentence and Monstream builds the stream for you.Create your own →