AbAg-XM · deep-N
Tenstorrent Tenstorrent Computed on Tenstorrent using TT-Bio

labelled samples · 4 models · the 164-target 2026ARK-AB benchmark

Sampling scales. Selection does not.

Draw 512 structures for one antibody–antigen complex and the best of them keeps getting better, all the way to the cap. The one the model hands you stops improving after about 32. Past that point every extra structure you pay for is one you will not receive.

Mean DockQ, 0 to 1 agreement with the solved structure, on the 160 targets all four models folded.

The vertical order of the solid lines is the ranking. The dashed lines rank differently: the model that hands you the least is not the one with the least to give.

Those intervals are marginal on the level and wide, because targets differ enormously. The paired 32 to 512 change is far tighter: delivered , best in the pool .

Best in the pool is the highest-DockQ structure among the k drawn, picked with the answer key: an upper bound no selector can beat, and what oracle-best-of-N means in this literature. Delivered is the one the model's own confidence ranks first, with no ground truth. On the DockQ scale 0.23 is the community's acceptable bar, 0.49 medium and 0.8 near-perfect.

Every point is an exact closed-form expectation over uniform k-subsets, never simulated. Intervals are 20,000 resamples over targets, one shared draw. The pooled 32 to 512 numbers average per target before bootstrapping, not four published intervals. This figure uses the 160 targets every model folded; the rest of the page uses each model's own 160 or 161, which moves a curve by at most .

What you actually get stops improving.

What confidence ranks first, out of k drawn, rises to about k = 32 and then goes flat. Over 32 draws to 512, .

Measured against each model's own single draw, so the axis reads as what drawing more bought. Band = 95% paired bootstrap on the delivered gain; the table is delivered(512) − delivered(k).

The ceiling keeps rising, so the gap keeps widening.

The ceiling gains DockQ per doubling, a rate unchanged from 64→128 to 256→512, so nothing saturates inside the measured range. The gap is DockQ wider at 512 samples than at 16, every interval excluding zero.

Correcting the shallower 256-sample round, which showed no knee: the second difference of the gap is negative with an interval excluding zero for , so the widening decelerates slightly there. Because the ceiling's gain per doubling is flat, that deceleration is the delivered line ticking up inside its own noise, not the ceiling bending.

Why: confidence stops ranking exactly where it is used.

The same score, correlated two ways. Across targets it works: a model knows which complexes it will fold well. Within one target's 512 samples it is near zero, ipTM included.

This run recorded only pTM and the ranking score for ESMFold2, so it has no ipTM row; its ranking score is mean pLDDT. For the other three the shipped selector is exactly their confidence_score.

Restrict the correlation between confidence and accuracy to the top tail a selector actually picks from and it collapses to . For OpenDDE-abag it inverts: inside its own top quartile, higher confidence means a worse structure. A random subset of the same size holds the whole-pool correlation, so that correlation is carried entirely by samples no selector will ever pick.

We tried to beat it. Nothing did.

Six alternative selectors, fixed in advance, each built only from numbers a model already returns. The bar, set before running them: beat the shipped selector with an interval excluding zero on two of the four models. Every one lands level with it or worse.

Combinations average each sample's rank under the listed scores. Structural consensus, ranking samples by agreement with the pool's modal pose, needs the structures themselves and is not tested here.

Effective N does not grow with N.

Effective N is the draws a perfect chooser would have needed to match what the model delivers from N. Across a 256-fold range it stays between , a fitted slope of against 0 for flat, so the share of your draws that reaches you is at N = 512.

What it would cost to get there by sampling alone.

Extrapolate the ceiling's trend to where 80% of targets carry an acceptable pose. Solid is MEASURED to 512; past the rule everything is DERIVED from a log-linear fit assuming no ceiling. Fitted counts: .

Spreading compute across models beats more samples, in principle.

Same card-hours per target, split four ways instead of poured into one model. From 0.08 card-hours the split beats every single model at the same spend, including the antibody specialist ( more targets solved, paired 95% interval excluding zero), and it overtakes the best single model run at 31× the compute from card-hours.

And the catch: nobody can harvest it yet

Picking the globally highest-confidence structure across all four pools gets worse as the budget grows.

Most failures never find the epitope.

The epitope is the patch on the antigen the antibody binds. On a target a model fails, its best sample overlaps the true epitope by a median of , against on one it solves. Two clean modes, almost nothing between: sampling refines poses far faster than it finds sites.

Look at one target.

One real 64-sample fold. Each mark is a structure, ordered by the model's own confidence. You get the leftmost.

How this was measured

What this is

Four independently trained predictors, each sampled to 512 structures per target on the 2026ARK-AB benchmark, every structure scored with DockQ against the experimental reference. Boltz-2, OpenDDE-abag and Protenix-v2 are AF3-style all-atom diffusion co-folders; ESMFold2 is a single-sequence folder. labelled samples, folded on Tenstorrent Wormhole Galaxy hardware.

They do not fail on the same targets: the pairwise overlap of their failure sets is (Jaccard), and are failed by all four. That partial independence is what the compute-split section buys.

Prior work

That deep sampling raises the ceiling faster than confidence ranking delivers is already published for antibody–antigen complexes, not new here. Fromm et al. (Bioinformatics 2026) report best-of-N mean DockQ rising from below 0.3 to above 0.5 by 200 samples across 110 targets, and name identifying the best generated model as the remaining bottleneck. OpenDDE's technical report measures the same ranked-versus-oracle gap at its default N = 5.

This panel adds that gap expressed as an effective N, the split of ranking failure into within-target and across-target, the finding that most failures are epitope-level rather than pose-level, the price of model diversity against sampling depth at measured compute, and a pre-declared null on fixing selection with the scores a model already returns.

Limitations, denominators and provenance

Reproduce

Full statistics, every interval and every control are in the findings document.

What this is built on

The result belongs as much to the people who released these models and tools. Every prediction here was made by someone else's model, scored by someone else's metric, against structures someone else solved and deposited.

All four models were run as shipped, on stock checkpoints, with each model's own default settings. No model was tuned, retrained or recalibrated for this benchmark.

How to cite

Thüning, M. (2026). AbAg-XM: 335,360 DockQ-labelled antibody-antigen structure predictions from four models. Tenstorrent. https://huggingface.co/datasets/Tenstorrent/abag-xm

@misc{abagxm2026,
  title  = {AbAg-XM: 335,360 DockQ-labelled antibody-antigen structure predictions
            from four models},
  author = {Th\"uning, Moritz},
  year   = {2026},
  url    = {https://huggingface.co/datasets/Tenstorrent/abag-xm},
  note   = {Analysis: https://moritztng.github.io/abag-scaling/}
}

The dataset and this analysis are CC-BY-4.0. The reference structures come from the PDB under CC0; if you use those, cite the original depositors too.