AIFlow Math Ink 0.6은 수학 필기를 이미지보다 point/stroke sequence로 먼저 처리하는 온디바이스 연구 모델이다. 원본 touch event를 보존하면서 모델 입력만 6Hz canonical tap으로 재표본화하고, 각 고립 기호의 378-class top-k와 visual-family 확률을 반환한다.
현재 공개본은 seed 17·31·47의 연구 checkpoint와 grouping/behavior head를 포함한다. 전체 수식 LaTeX decoder나 완성된 Android LiteRT 배포본은 아니다.
AIFlow Math Ink 0.6 architecture
2026-07-24 federation 감사
출처 이름이나 sample ID만 세지 않고, 이동·크기를 제거한 trajectory signature와 writer split을 다시 감사했다.
항목
교정 전
교정 후
UJI Pen v1↔v2 중복
1,364
0
HWRT writer split 누수
93명
0
전체 origin overlap
0
0
유효 supervised source
4
3
UJI Pen v1은 v2에 전부 포함된 mirror라 sampler에서 제거했다. HWRT 공식 test는 제품 경로에서 제외하고, 승인 train writer 275명을 별도의 train/validation/test로 다시 나눴다. UJI v2는 공식 test를 보존하면서 train writer validation을 격리했다.
현재 registry는 조사 40/200, 승인 배포 독립 그룹 7/30이다. 교정 후 seed-17 federation checkpoint에는 실제 training_source_ids와 독립 source group, registry SHA-256이 기록된다. 다만 30개 독립 승인 그룹·3-seed·P writer/device-disjoint release gate가 아직 없으므로 공개 모델은 계속 연구용이며 product_validation=false다. 교정 전 federation 정확도는 제품 근거로 사용할 수 없다.
교정 후 seed-17은 checkpoint에 실제 source ID·독립 그룹·registry SHA-256을 기록한다. UJI 전체 train을 포함한 15,829개 source-balanced 학습에서 test online top-1은 Pendigits/UJI v2/HWRT가 93.2% / 71.8% / 84.0%, raster top-1은 42.8% / 41.0% / 74.4%다. GTX 1650에서는 batch 32가 안정적이고 batch 64는 virtual top-4 peak OOM이었다.
이 결과는 출처 cap을 풀어도 UJI writer-generalization과 exact case family가 남는다는 근거다. Raster encoder·decoder만 여는 shadow run은 Pendigits를 개선했지만 HWRT holdout을 손상해 자동 guard에서 거부됐고 배포 모델에 포함하지 않는다. 3-seed distillation은 현재 보류 상태다.
Online affine(회전 ±10°·scale ±8%) full-source 1-epoch ablation도 validation macro online top-1 88.57% → 88.07%로 하락했고, UJI는 80.00%로 그대로인 반면 HWRT online/raster top-1은 90.04/78.16% → 88.89/76.63%가 됐다. 선택 결과는 기준선 epoch 0/alpha 0.0이므로 이 증강은 채택하지 않는다.
긴 GPU 실험에는 epoch-state를 남긴다. state에는 student/best/anchor model, Adam, source-balanced sampler generator, Python/NumPy/PyTorch/CUDA RNG이 포함되며 base SHA-256·seed·source·split contract가 모두 일치할 때만 resume한다.
동일 writer의 고립 glyph를 baseline anchor로 쓰는 size-resolver proxy도 UJI에서 51.09% → 45.63%, HWRT에서 53.57% → 39.29%로 악화됐다. 이는 writer 묶음이 실제 수식 행 문맥이 아니라는 반증이다. 따라서 relative-size resolver는 현재 isolated symbol output에는 적용하지 않고, P Formula의 실제 행/Tray context가 확보된 경우에만 다시 검증한다.
2026-07-25: 기각한 generalization loop와 재현성 보강
기하 일관 affine+raster-loss 후보는 좌표 변환 뒤 shape/curvature/bbox/aspect/speed도 다시 계산하도록 수정해 재검증했지만, source macro online이 88.57% → 88.07%, HWRT online/raster가 90.04/78.16% → 88.89/76.63%로 하락해 채택하지 않았다.
이 반복에서 과거 stage-3 checkpoint가 현재 loader의 training selection으로 정확히 재현되지 않는 provenance 공백도 확인했다. 이후 trainer는 source별 선택 수와 (source, sample_id, origin_id) SHA-256을 checkpoint·report·epoch-resume contract에 넣어, 같은 표본 수이지만 다른 원본을 섞는 재개를 fail-closed로 막는다. product_validation은 계속 false다.
Provenance baseline v1
새 seed-17 기준선은 정확히 기록된 15,829 training record(Pendigits 4,457 / UJI v2 5,372 / HWRT 6,000)와 sample-origin SHA-256 fc103edba331d3682875171b04942a2028232a401a0ca72cd0b1c52da5608c91를 사용한다. Held-out online top-1/top-5는 Pendigits 95.0/99.4%, UJI 73.2/94.6%, HWRT 84.2/99.0%다. 전체 7,031개 online은 86.06/97.50%, visual-family 90.91%, writer p10 60.0%다. 92/99 online, 90% raster, 30 independent P-source, 3-seed gate에는 모두 미달하므로 연구 기준선일 뿐 배포 모델이 아니다.
Checkpoint 내부 경로는 lineage 기록이며 로컬 절대경로에 의존해 추론하지 않는다.
현재 성능
378-label trajectory classifier
최신 online case-context seed 17 기준:
split
top-1
top-5
family top-1
writer validation, 4,261 samples
86.13%
99.48%
92.94%
paired writer-disjoint test, 3,782 samples
82.87%
97.73%
91.22%
Paired source는 HWRT 내부 writer hash split과 UJI Pen v1/v2 writer-disjoint 계약을 사용했다. HWRT official test는 거대 writer 중복 때문에 제외했다.
수식 행동 head
정답 symbol grouping 이후의 CROHME 조건부 역할 평가:
지표
3-seed 평균
role accuracy
93.33%
macro-F1
76.41%
lowercase identifier recall
95.66%
uppercase identifier recall
55.91%
multiplication recall
88.89%
ECE
4.88%
행동 head를 실제 teacher 출력 뒤에 연결하고 validation에서만 confidence threshold를 선택한 결과, CROHME official test의 대상 기호 exact top-1은 33.65%에서 47.30%로 평균 13.65%p 상승했다. Rewrite precision은 81.35%였다. 다만 같은 구간의 teacher visual-family top-1이 53.02%에 불과해, 행동 문맥만으로 잘못된 형태군을 복구할 수 없었다.
R-track 연속 수식 어댑터
제품 trajectory encoder와 online adapter는 고정하고, CROHME 연속식의 정답 symbol group 위에서 92KB zero-init formula adapter만 학습했다. 이 checkpoint는 구조 검증용 R_noncommercial_only 모델이며 제품 weight 또는 distillation 입력이 아니다.
3-seed 지표
평균
최저
writer-validation exact top-1
85.19%
82.91%
writer-validation visual-family top-1
93.30%
93.03%
official test exact top-1
82.38%
81.13%
official test visual-family top-1
88.17%
88.03%
visual-family 일반화 gap
5.13%p
—
Validation에서는 세 seed 모두 형태군 92%를 넘겨 현재 TCN 구조가 연속 수식에도 적응할 수 있음을 확인했다. 반면 unseen official test에서는 모두 실패했다. 현 병목은 모델 용량보다 writer/source domain 일반화이며, 다음 gate는 상용 허용 P-track 연속식의 writer/device-disjoint 재학습이다.
P 고립기호 합성 수식 proxy — 기각
승인 HWRT/UJI trajectory 1,590개를 숫자 anchor 사이에 합성 배치해 행동 head에 weight 0.35로 추가한 seed-17 실험은 accuracy가 93.07%로 같았지만 macro-F1 76.64→73.62%, uppercase recall 61.29→45.16%로 악화됐다. 실제 연속식 validation과 비교한 formula-relative bbox width/height의 최대 절대 SMD는 2.82로 호환 기준 0.5를 크게 넘었다.
고립기호에는 전체 formula bbox·Tray·이웃 부재 분포가 없으므로 임의 합성 배치를 행동/formula 제품 학습에 사용하지 않는다. 이 경로는 rejected_no_3seed_expansion이며 실패 checkpoint도 배포하지 않는다. 승인 고립기호는 shape encoder에만 유지하고, 행동 학습은 실제 AIFlow P Formula v1 연속식을 기다린다.
Grouping boundary head
지표
이전
boundary 적용
CROHME exact partition
60.04%
60.25%
pair-F1
91.07%
91.26%
overmerge formula rate
21.72%
20.49%
P boundary shared-state 정정
초기 공개 P delta는 online_adapter.pt의 shared_state_dict를 적용하지 않은 loader 결함을 상속했다. 이 때문에 강한 main encoder를 빠뜨린 낮은 기준선과 비교했으며, 해당 auxiliary/joint checkpoint와 성능 주장을 철회했다.
올바른 합성 순서인 base → adapter shared state → modality adapter → optional head로 세 seed를 다시 학습한 결과는 다음과 같다.
paired proxy test
정정된 main baseline
joint 결과
변화
exact top-1
83.03%
82.81%
-0.22%p
family top-1
91.37%
90.88%
-0.49%p
single-symbol recall
—
92.83%
95% floor 실패
cross-boundary recall
—
98.47%
통과
세 seed 모두 release gate를 실패했으므로 joint delta는 배포하지 않는다. 현재 유효한 main 구성은 seed별 base_378.pt + online_adapter.pt이며 adapter 안의 shared_state_dict를 반드시 먼저 적용해야 한다.
Joint delta 없는 main+shadow auxiliary device stress에서 최악 exact/family 하락은 affine 변형의 -0.98%p/-1.08%p였다. Software stress는 통과했지만 clean single recall이 91.42~94.50%이므로 auxiliary head도 제품 채택 대상이 아니다.
Composite torch.export
Export 그래프는 이제 base-only가 아니라 online/raster modality adapter와 shared state를 포함한다. 세 seed 모두 각 경로 대표 입력 76개에서 eager 대비 top-1 100% 일치, 최대 logit 절대오차 0.0을 기록했다. Validation-only calibration으로 고정된 online family-fusion 0.15도 export에 포함된다. Fresh seed-17 online/raster .pt2 합계는 24,300,116 bytes다.
.pt2는 Android용 .tflite가 아니다. LiteRT Torch 0.9.1 변환과 Android runtime parity는 아직 완료되지 않았으므로 litert_exported=false를 유지한다.
CPU latency·memory proxy
Family-fusion 0.15를 포함한 Windows PyTorch CPU 재측정에서 online p95는 8.909.67ms, raster p95는 22.1026.63ms였다. Tensor state는 8.95MB, 모델 로드 후 inference RSS 증가분을 합친 구조 proxy는 26.39MB다. 전체 Python process RSS는 PyTorch runtime을 포함하므로 Android LiteRT memory 근거가 아니다.
Online 오류 합의 감사
정정 composite 세 seed를 선택에 쓰지 않은 paired writer/device-disjoint 3,782개에서 다시 감사했다.
지표
결과
exact ensemble top-1
83.71%
exact ensemble top-5
98.02%
visual-family top-1
92.99%
seed oracle top-1
88.05%
세 seed 공통 오류
11.95%
Exact 오류의 56.98%(전체 9.28%p)는 O/0/o, 대소문자, 수직선·cross처럼 같은 visual family 안의 의미 혼동이다. 따라서 0.6의 trajectory 단계는 형태군 후보를 반환하고, exact 의미는 실제 수식 행의 상대 크기와 행동 문맥이 결정해야 한다. 실제 P 연속식이 없으므로 formula-context exact 92% gate는 아직 통과하지 않았다.
Validation에서 family-fusion 0.15를 선택한 뒤 paired-test에 한 번 적용하자 exact top-1은 83.71→83.82%(+0.11%p)였다. 이는 작은 보정이며 문맥 레이어를 대체하지 않는다.
이미지, raw stroke, virtual stroke는 서버 payload로 전송하지 않는다. virtual hypotheses는 로컬 debug에서만 노출한다.
Local P Formula intake
Drawer의 local-only AIFlow Ink v1은 자동으로 학습 정답이 되지 않는다. 별도 human annotation JSONL이 실제 비식별 device_id, label_status=human_verified, 모든 formula_cell_id → token 전단사를 제공해야 한다.
Materializer는 다음 조건을 모두 통과한 뒤에만 UTF-8 P Formula v1 JSONL을 원자적으로 생성한다.
모든 raw stroke가 정확히 한 symbol group에 포함됨
timestamp·pressure 결측을 원본 그대로 보존
기준 checkpoint의 중복 없는 378 exact vocabulary와 token 일치
실패 시 report만 남고 dataset은 생성되지 않는다. 이 기능은 서버 전송이나 자동 수집을 수행하지 않는다.
P-only formula adapter training
Materialize된 P Formula v1은 전체 formula bbox 기준 128×19 symbol tensor로 변환된다. Product encoder와 online adapter는 동결하고 hidden-64 zero-init formula adapter만 학습한다.
text
1family CE + exact CE×0.10
2context dropout 0.30
3inverse-sqrt(source frequency × exact-label frequency) sampler
4validation-only checkpoint selection
각 seed는 test exact top-1 92%, top-5 99%, macro-F1 90%, writer floor 75%, 결측 metadata slice 하락 3%p 이하를 모두 통과해야 한다. Seed 17·31·47이 개별 통과하고 세 run의 원본 P Formula JSONL SHA-256이 정확히 같은 경우에만 single mobile student distillation을 허용한다. Teacher ensemble 자체는 기기에 탑재하지 않는다.
실제 seed-17 composite와 GTX 1650을 사용한 1-epoch fixture smoke에서 CUDA 학습부터 report/checkpoint 생성까지 통과했다. Fixture checkpoint는 성능 근거가 아니므로 이 공개 저장소에 포함하지 않았다. 고정 recipe는 configs/MATH-INK-06-P-FORMULA-v1.json에 있다.
Single-student distillation
통과한 세 teacher는 base → shared online adapter → P formula adapter 순서로 합성한다. Temperature 2.0의 exact/family 확률을 seed 사이에서 평균하고 KL + hard-label CE로 hidden-64 formula adapter 하나만 학습한다. Student checkpoint에는 teacher weight를 포함하지 않는다.
Student는 자체 92/99·macro-F1·writer/missing gate뿐 아니라 teacher ensemble 대비 exact top-1·top-5·visual-family 하락 1%p 이하를 모두 만족해야 한다. 2/2/2-symbol fixture의 3-teacher→student CUDA 실행 경로는 통과했지만 정식 student gate는 실패했다. Fixture와 checkpoint는 이 공개 저장소에 없으며, 실제 P 데이터 성능이나 제품 검증을 뜻하지 않는다. LiteRT 변환도 아직 수행하지 않았다.
Fail-closed release and export
run_math_ink_06_p_formula_release.py는 seed 17/31/47 학습·개별 gate, 동일 data SHA-256 요약, single-student 증류, student export를 순서대로 실행한다. 한 단계의 정식 gate가 실패하면 다음 단계는 실행하지 않는다.
Student export는 online adapter → formula adapter → shared classifier → family fusion 순서를 graph에 포함한다. P track·세 teacher lineage·동일 corpus·teacher weight 미포함을 다시 확인하고, 실제 P test representative에서 strict eager/export top-1 100%, logit 최대 오차 0.02 이하, graph 25MB 이하를 요구한다. 실패 fixture는 정식 export에서 거부된다. 378-label 메모리 graph의 strict export 호환성은 top-1 2/2·오차 0으로 확인했지만 실제 LiteRT/Android 검증은 아니다.
Android ODA runtime
android/aiflow-math-ink-runtime은 recognizeOnline, recognizeRaster, SymbolResult와 AIFlow Ink v2의 raw stroke·6Hz tap·128×19 계약을 구현한다. Server payload에는 후보·confidence·model version·latency만 포함된다.
Android canonicalizer는 Python 대표 4행×19채널과 절대오차 1e-4 이내 parity를 통과했다. 최신 standalone LiteRT CompiledModel 2.1.6으로 APK asset의 online/raster graph를 직접 실행하며 ML Kit custom model API를 사용하지 않는다.
Raster 배포 graph는 exact logits뿐 아니라 top-4 coordinates/state logits/progress/hypothesis scores를 포함한 정확히 5개 output을 요구한다. recognizeRasterDebug는 이 값을 로컬에만 노출하고 server payload에는 포함하지 않는다. 실제 seed-17 378-label·76 representative strict export에서 online/raster top-1 76/76, 최대 오차 0.0과 다섯 output shape를 확인했다.
이 저장소의 Colab ZIP은 공개 가능한 R_public_conversion_only 연구 graph다. P student와 P corpus를 포함하지 않으며 product bundle이 아니다. 제품용 builder는 별도로 제공하지만 생성물에는 P data가 들어가므로 public upload와 server 전송이 금지된다. 새 P student는 base·online adapter SHA-256까지 checkpoint lineage로 가져야 한다.
3-tier benchmark runner는 동일 online/raster SHA-256 쌍의 representative로 warm-up 후 각 100회, PSS, battery charge-counter delta를 측정한다. 개별 hash는 순서가 고정된 bundle SHA-256으로도 묶는다. Python 요약기는 low/mid/high가 정확히 하나씩이고 version·개별 hash·bundle hash가 동일할 때만 online p95 50ms, raster p95 200ms, PSS 100MiB, battery 계측 gate를 AND로 결합한다. 최종 product manifest는 P 3-seed, 배포 raster 품질, LiteRT model pair, Android 3-tier의 동일 lineage를 다시 확인한다. Kotlin unit test 8개와 전체 Python 회귀 355개, release AAR build가 통과했다. AAR은 84,436 bytes, SHA-256 2c65d4af59af1204fad66eff15bec5ef125dfa242da82d05b0e85cf99504c44d이며 모델을 포함하지 않는다. 실제 flatbuffer·기기 benchmark 값은 아직 없다.
알려진 한계
378-label paired writer/device-disjoint top-1 목표 92%에 아직 미달한다.
전체 수식 LaTeX, Tray decoder, gridding은 0.7 범위다.
uppercase 역할 recall과 O/0, styled-letter hard family가 남은 병목이다.
행동 exact gate는 유효하지만, 연속식 teacher 형태군이 틀리면 역할 head가 복구할 수 없다.
R-track formula adapter는 validation 형태군 93.30%를 달성했으나 official test 88.17%에 그쳐 제품 검증을 통과하지 못했다.
승인 P 고립기호 합성 행동 proxy는 layout SMD 최대 2.82와 uppercase recall 하락 때문에 기각했다.