A contrastive encoder that aligns HAADF-STEM microscopy images with their acquisition metadata in a shared 128-d embedding space. This variant uses a ResNet-18 image encoder trained from scratch on 256x256 patches at effective batch size 512, with a 3-layer MLP metadata encoder of width 256.
Model Details
Architecture: ResNet-18 image encoder (trained from scratch, single-channel input) + 3-layer MLP metadata encoder (hidden dim 256)
A Ridge regression ($\alpha = 1.0$) trained on the frozen visual embedding recovers all seven acquisition parameters. Coefficient of determination ($R^2$), SMAPE (in physical units), and Pearson $r$:
Dimension
$R^2$
SMAPE
Pearson $r$
pixel_size
0.720
41.8%
0.849
dwell_time
0.782
28.4%
0.885
convergence_angle
0.537
13.0%
0.733
beam_current
0.655
34.7%
0.811
gain
0.841
5.3%
0.917
offset
0.802
9.2%
0.898
inner_coll_angle
0.499
9.5%
0.708
Mean
0.691
20.3%
0.829
The higher SMAPE on pixel_size, dwell_time, and beam_current is expected: those dimensions are stored log10-transformed because they span several orders of magnitude in physical units, so small residuals in log-space amplify when exponentiated back.
@misc{cimp2026,
title={Contrastive Image-Metadata Pre-training for Materials Transmission Electron Microscopy},
author={Channing, Georgia and Keller, Debora and Rossell, Marta D. and Torr, Philip and Erni, Rolf and Helveg, Stig and Eliasson, Henrik},
year={2026},
}