Skip to content

Dashboard

Benchmark status

The Perturb-LM validation ladder. Only baseline rows with completed controls show numeric scores; frozen and projected encoder rows remain pending until real runs are executed and reviewed.

Illustrative interface demo — not real model output
Illustrative interface demo — not real model output. This dashboard reflects benchmark scaffolding and completed baseline controls. No BiomedBERT, linear-projection, replicate-consensus, or held-out-batch model scores are shown.

Current status

Where the project stands

Foundation

Pipeline scaffold, preprocessing, query construction, and public-copy consistency are validated.

Validated

Lexical benchmark

Identifier-stripped TF-IDF baseline completed with query-bootstrap CI.

Validated

Real encoder run

BiomedBERT is selected and pinned; frozen embeddings have not produced benchmark results.

Pending

Linear projection result

Regularized text-to-profile projection results remain pending.

Pending

Held-out batch

The current local dataset has one inferred batch; batch generalization cannot be claimed.

Unavailable

Live model integration

No live model is served by the public prototype.

Unavailable

Benchmark overview

Public-safe aggregates only

Profiles
4,524
Features
904
Queries
641
TF-IDF mAP
0.2513

identifier-stripped TF-IDF

95% CI
0.2445 to 0.2582

95% query-bootstrap

Model status
pending

No learned model result is public.

Selected encoder
BiomedBERT

Revision pinned; run pending.

Method comparison

Pending rows do not render scores

Charts and tables use the same public-safe source. A method with no completed result shows a pending badge and no fabricated bar.

Method comparison table. Pending learned-model rows have no numeric score or bar.
MethodmAP95% CIStatus note
Random
0.0034
not availableUniform random ranking control from the corrected full-query baseline.
Shuffled labels
0.0018
not availableLabel-permutation control preserving benchmark marginals.
Identifier-stripped TF-IDF
0.2513
0.2445 to 0.2582Primary lexical baseline over identifier-stripped query text.
Frozen BiomedBERT embeddingsPendingnot availablePending real experiment. No score is shown until the run completes.
Frozen embeddings + linear projectionPendingnot availablePending real experiment. Requires frozen embeddings and train-only fitting.
Replicate-consensus alignmentPendingnot availablePlanned ablation. No consensus-model score exists yet.
Held-out batch evaluationPendingnot availableUnavailable until a second compatible batch is integrated.

Text equivalent: identifier-stripped TF-IDF is the primary baseline at 0.2513 mAP. Learned-model rows are pending and intentionally have no numeric score.

Split and filter readiness

Each condition must report its own query counts

The dashboard shows configured readiness only. It does not reuse the full 641-query count as if every split-specific model run had completed.

Held-out plate

Configured for split-specific evaluation with independently reported query counts.

Ready

Held-out treatment

Configured for unseen-treatment evaluation without train/test treatment overlap.

Ready

Exclude same plate

Retrieval filter removes same-plate candidates.

Ready

Exclude same well

Retrieval filter removes same-well candidates.

Ready

Exclude same plate and well

Strict combined filter for leakage-aware ranking.

Ready

Held-out batch

Requires a second compatible batch.

Unavailable

Leakage controls

Controls enforced before stronger claims

  • Treatment identifiers are removed from identifier-stripped text.
  • Target sequences are prohibited from query construction.
  • Plate, well, and batch identifiers are excluded from public query text.
  • Preprocessing and model fitting must use training data only.
  • Every split/filter must report its own evaluable-query count.
  • Public reports expose aggregate metrics only, not row-level CPJUMP1 records.

Phase 3C roadmap

Next experiments, still gated

  1. 01

    Frozen BiomedBERT baseline

    Next

    Run frozen embeddings against the benchmark and report leakage-aware metrics.

  2. 02

    Regularized linear projection

    Next

    Fit a small projection with train-only preprocessing and held-out evaluation.

  3. 03

    Replicate-consensus profiles

    Planned

    Compare individual profile targets to mean or median perturbation consensus targets.

  4. 04

    Hard-negative-ready evaluation

    Planned

    Prepare biologically related but label-distinct comparison groups.

  5. 05

    Second-batch harmonization

    Planned

    Integrate a compatible batch before making batch-generalization claims.

  6. 06

    Representative image linkage

    Exploratory

    Future qualitative inspection only; current benchmark remains profile-based.

Limitations

What this dashboard does not claim

  • The real biomedical text-alignment experiment has not started.
  • Split-specific learned-model results remain pending.
  • Held-out-batch evaluation is unavailable.
  • The additional profile plate changes the feature schema and is only a compatibility investigation.
  • The current benchmark is text-to-morphology-profile retrieval, not validated text-to-image retrieval.