Kreasof AI
← Back to the article

ATMA Ablation Archive

Filter every 1B-token recipe-selection run, rank metrics at each evaluation length, inspect resolved configurations, and compare validation trajectories.

This archive supports a research article currently under review. Stage I selects a recipe across nuisance-factor settings; it is not a seed study or a substitute for the larger matched comparison.

120 / 120factorial cells complete
+47.8mean Polar memory effect at 64K
20 / 20matched Polar cells improved
92.5%selected Stage I cell at 64K

Explore the factorial

Choose a metric, filter the recipe axes, then select rows to compare validation trajectories.

Validation trajectory

The top three ranked cells start selected. Use the Compare column below to add or remove trajectories.

Ranked cells

Click a row to inspect its metrics and fully resolved configuration.

The promoted recipe at 9.816B tokens

The matched attention variants use memory, no local window, no distractor alignment, length-2K training, and the same L40S environment. Raven models are separately optimized comparison points.

GroupModelToken @2KToken @256KExact @256KBABILong @256KBPB ratioBase mean
MatchedNoPE99.2%0.6%0.0%0%5.65×45.15%
MatchedPolar98.9%34.4%9.0%28%1.26×43.29%
MatchedRoPE88.2%0.0%0.0%14%1.75×44.60%
Separately optimized recurrent family
Raven / AdamWAtma-Raven-Titans43.1%10.0%0.0%38%1.02×43.05%
Raven / AdamWRaven Native41.1%20.3%0.0%46%1.01×43.05%

Retrieval averages tasks, haystack suites, and needle depths. Exact requires all five target tokens under teacher forcing. BPB averages three fixed-target datasets. Base mean averages eight 2K zero-shot tasks.

One extreme retention head mediates part of the extrapolation failure

A parameter scan found maximum zero-input half-lives of 21.01M tokens in NoPE, 3.07M in Polar, and 682 in RoPE. We capped only the largest runtime head at a 256-token half-life without rewriting checkpoint weights, then reran downstream, retrieval, fixed-target BPB, and BABILong.

Length-wise untouched and capped curves for retrieval, BABILong, and long-document bits per byte.
The recovery is length-wide rather than an endpoint artifact. NoPE BABILong changes from 21/6/0/0/0% to 54/45/37/35/39% at 16/32/64/128/256K, while Raven Native and Atma-Raven-Titans remain visible as untouched recurrent-family references.
ModelMax half-lifeBase meanBPB @256KRetrieval @256K
token / exact
BABILong @256K
NoPE21.01M45.15 → 45.108.097 → 1.5950.60 / 0.00 → 0.93 / 0.000 → 39
Polar3.07M43.29 → 43.251.825 → 1.50134.37 / 9.00 → 47.07 / 16.3328 → 42
RoPE68244.60 → 44.662.512 → 2.4600.00 / 0.00 → 0.00 / 0.0014 → 11

Baseline → capped. The full rerun compares clamped jobs with pinned archived baselines; a separate same-process paired sweep establishes the intervention. NoPE exact retrieval remains zero and RoPE does not recover, so the cap is a selective mechanism probe rather than a universal improvement. The paper appendix reports every downstream task and every retrieval, BPB, and BABILong length cell.

Metric definitions and composite scores

Length-weighted metrics use weight L / 2048, emphasizing extrapolation endpoints. Composite indices are exploratory archive aids and are not article headline metrics.