Foundation
Pipeline scaffold, preprocessing, query construction, and public-copy consistency are validated.
Dashboard
The Perturb-LM validation ladder. Only baseline rows with completed controls show numeric scores; frozen and projected encoder rows remain pending until real runs are executed and reviewed.
Current status
Pipeline scaffold, preprocessing, query construction, and public-copy consistency are validated.
Identifier-stripped TF-IDF baseline completed with query-bootstrap CI.
BiomedBERT is selected and pinned; frozen embeddings have not produced benchmark results.
Regularized text-to-profile projection results remain pending.
The current local dataset has one inferred batch; batch generalization cannot be claimed.
No live model is served by the public prototype.
Benchmark overview
identifier-stripped TF-IDF
95% query-bootstrap
No learned model result is public.
Revision pinned; run pending.
Method comparison
Charts and tables use the same public-safe source. A method with no completed result shows a pending badge and no fabricated bar.
| Method | mAP | 95% CI | Status note |
|---|---|---|---|
| Random | 0.0034 | not available | Uniform random ranking control from the corrected full-query baseline. |
| Shuffled labels | 0.0018 | not available | Label-permutation control preserving benchmark marginals. |
| Identifier-stripped TF-IDF | 0.2513 | 0.2445 to 0.2582 | Primary lexical baseline over identifier-stripped query text. |
| Frozen BiomedBERT embeddings | Pending | not available | Pending real experiment. No score is shown until the run completes. |
| Frozen embeddings + linear projection | Pending | not available | Pending real experiment. Requires frozen embeddings and train-only fitting. |
| Replicate-consensus alignment | Pending | not available | Planned ablation. No consensus-model score exists yet. |
| Held-out batch evaluation | Pending | not available | Unavailable until a second compatible batch is integrated. |
Text equivalent: identifier-stripped TF-IDF is the primary baseline at 0.2513 mAP. Learned-model rows are pending and intentionally have no numeric score.
Split and filter readiness
The dashboard shows configured readiness only. It does not reuse the full 641-query count as if every split-specific model run had completed.
Configured for split-specific evaluation with independently reported query counts.
Configured for unseen-treatment evaluation without train/test treatment overlap.
Retrieval filter removes same-plate candidates.
Retrieval filter removes same-well candidates.
Strict combined filter for leakage-aware ranking.
Requires a second compatible batch.
Leakage controls
Phase 3C roadmap
Run frozen embeddings against the benchmark and report leakage-aware metrics.
Fit a small projection with train-only preprocessing and held-out evaluation.
Compare individual profile targets to mean or median perturbation consensus targets.
Prepare biologically related but label-distinct comparison groups.
Integrate a compatible batch before making batch-generalization claims.
Future qualitative inspection only; current benchmark remains profile-based.
Limitations