Aggregate representations did not beat simple baselines.
Firing-rate and network summaries failed to improve experiment-level performance prediction under organoid-held-out evaluation. A rank association did not become calibrated predictive accuracy.