MicroT-test1-500K-UltraChat is a 498,528-parameter vanilla decoder-only transformer — multi-head causal self-attention with RoPE — pretrained on UltraChat 200k conversation data (instead of the project's Discord-Dialogues baseline) and then fine-tuned with FMSP on 9,012 general-knowledge QA pairs.
This is the 500K member of the MicroT-test1 family: the registered attention-based reference baseline of the MicroMixer-4 project, here rerun on UltraChat 200k as part of the dataset-efficiency comparison study — six parameter budgets × two architectures × two open pretraining corpora, all fine-tuned with the identical P05 FMSP recipe at seed 42. Analysis.
It is deliberately boring — the standard 2018–2020 transformer recipe parameterized down to the sub-1M regime: no flash attention, no SwiGLU, no ALiBi, no QKNorm, no sliding window, no MQA/GQA, no MoE.
🏗️ Architecture
mermaid
1graph TD
2 A[Byte Input]--> B[Embed 256→96]3 B --> C[Transformer Block × 4]4 C --> D[RMSNorm]5 D --> E[LM Head Tied with Embed]6 E --> F[Byte Output]78subgraph"Transformer Block (pre-norm)"9 X[Input 96]--> N1[RMSNorm]10 N1 --> AT["MHA 6 heads × d_head 16<br/>RoPE θ=10000 on q,k · causal SDPA"]11 AT --> R1[+ residual]12 R1 --> N2[RMSNorm]13 N2 --> MLP["GELU MLP 96→424→96"]14 MLP --> R2[+ residual]15end1617style A fill:#007BFF,color:#fff18style F fill:#00D620,color:#fff19style AT fill:#FF6600,color:#fff
Model Configuration
Parameter
Value
Hidden Dimension (d_model)
96
Attention Heads
6 (d_head = 16 at every size)
Number of Blocks
4
FFN Hidden
424
Position Encoding
RoPE θ=10000 on q/k only (non-persistent buffers)
Attention
Causal MHA via F.scaled_dot_product_attention(is_causal=True)
Activation
GELU
Biases
None — no bias parameters anywhere
Normalization
RMSNorm (pre-norm)
Max Sequence Length
1024
Vocabulary Size
256 (byte-level)
Output Head
Tied with input embedding
Core Components
┌──────────────────────────────────────────────┐
│ Transformer Block (×4) │
│ h = h + MHA(RMSNorm(h)) # RoPE q/k, causal│
│ h = h + MLP(RMSNorm(h)) # GELU d→ffn→d │
│ no biases, no flash, no tricks — vanilla │
└──────────────────────────────────────────────┘
The d_head=16 contract is hard-asserted across all six sizes so that attention-head behavior is comparable at every budget and never confounds the memorization measurements.
🎯 Generation Examples
Questions the model was trained on (FMSP train set, 9,012 QA pairs — greedy decoding, repetition_penalty=1.2, no_repeat_ngram_size=4):
[Prompt] User: Who painted the Mona Lisa?
Assistant:
[Output] Leonardo da Vinci painted the Mona Lisa during the early 1500s in Italy.
correct — near-verbatim recitation of the trained fact
[Prompt] User: Who painted The Starry Night?
Assistant:
[Output] Vincent van Gogh painted The Starry Night in 1889 while staying at a mental asylum in Saint-Rém…
correct — recites the trained answer to the truncation limit
Questions the model has never seen and cannot answer (unanswerable probe — the correct behavior is to decline; the model's actual behavior is shown):
[Prompt] User: Who painted the Glimmering Frostberry?
Assistant:
[Output] The Gluta French Friedrich painted a kernel of Ariston that splaces through the marinade and lo…
fabricates — plausible-sounding nonsense on a nonexistent subject
[Prompt] User: Who composed the Symphony of Hollow Dawn?
Assistant:
[Output] Ludwig van Beethoven composed his Symphony of Hind on Beijing Teller.
fabricates — retrieves an unrelated trained fact instead of abstaining
498,528 parameters; capacity is the binding constraint on every axis
Reference baseline, not a product
Exists to score the Mixer against attention at matched budget
Knows only what it memorized
Knowledge is limited to the 9,012 trained QA pairs + the pretraining-corpus distribution
Does not abstain
Unknown questions are answered with fabrication or degeneration, not refusal — see the examples above
Byte-level noise
256-vocab byte tokenizer; PPL not comparable to BPE baselines
Research use only
Architecture/scaling research artifact, not a production model
🧬 Context
This is the 500K UltraChat-pretrained arm of the dataset-efficiency comparison study (24 arms: {V87, V88} × {UltraChat, SmolTalk2} × six sizes) in the MicroMixer-4 project. Each arm swaps the Discord-Dialogues pretraining baseline for an open corpus (UltraChat 200k) and reruns the identical P05 FMSP recipe at seed 42. Sibling repos: llaa33219/MicroMixer-4-{1M..10K}-{UltraChat,SmolTalk2} and llaa33219/MicroT-test1-{1M..10K}-{UltraChat,SmolTalk2}; the discord-pretrained baselines are llaa33219/MicroMixer-4-{1M..10K} and llaa33219/MicroT-test1-{1M..10K}. Full analysis: DATASET_COMPARISON_ANALYSIS.md.