Multi-Xi PARFLM (Property-Attractive-Repulsive Force Language Model)
The Multi-Xi PARFLM extends the SPLM with pairwise token interaction forces derived from a second scalar potential \(V_\phi\). While the base SPLM's \(V_\theta\) provides a single-body potential (each token interacts only with a summary of its past), PARFLM adds explicit pairwise forces \(V_\phi(h_t, h_s)\) between tokens -- the physics-informed analogue of attention's pairwise dot-product, but derived from a gradient of a scalar potential (making it conservative).
The pairwise forces use Gumbel-softmax top-k sparse routing to keep the cost at \(O(Tk)\) rather than \(O(T^2)\). This model achieves 12.06 PPL on TinyStories, a 2.6 PPL improvement over the standalone Multi-Xi SPLM.
Multi-channel K-EMA \(\xi\) (from the Multi-Xi SPLM): K=8 learnable causal exponential moving averages giving \(V_\theta\) a multi-resolution summary of the past.
Sparse PARF pair-interactions: A second scalar potential \(V_\phi(h_t, h_s)\) adds particle-exchange forces between token pairs, routed via Gumbel-softmax top-k selection.
Input tokens x_1, ..., x_T
|
Embedding E[x] + positional encoding
|
For each of L=8 integration steps:
|
+-- K-EMA channels: xi^(k)_t = causal_ema(h, alpha_k) [K=8 channels]
|
+-- Single-body: V_theta([xi_1..xi_K, h]) -> R [3-layer MLP]
|
+-- Pair routing: score_head(h_t, h_s) -> top-k selection [Gumbel-softmax]
|
+-- Pair forces: V_phi(h_t, h_s) -> R [structural competitive]
|
+-- Total: U_t = V_theta + sum V_phi
|
+-- Conservative force: f = -grad_h U_t [autograd]
|
+-- Damped Euler step: v += dt*f/m; v /= (1 + dt*gamma); h += dt*v
|
+-- LayerNorm(h)
|
Logits = h @ E^T [tied embeddings]
Parameter
Value
Hidden dim (d)
256
Layers (L)
8
\(V_\theta\) hidden / depth
1024 / 3
Xi channels (K)
8
Alpha init
log-spaced
\(V_\phi\) kind
structural_competitive
\(V_\phi\) hidden (H)
128
Sparse routing top_k
8
Gumbel tau
1.0 -> 0.1 (annealed)
Mass model
logfreq (frozen surprisal lookup)
Damping \(\gamma\)
0.30 nominal, ~0.033 effective (LayerNorm prevents compounding; see note below)
Total parameters
17,632,215
Effective damping. The nominal \(\gamma = 0.30\) overstates the true dissipation. The LayerNorm applied after each integration step rescales the hidden state, absorbing most of the velocity decay. The dynamics are therefore heavily underdamped even at this nominal value. Gamma-sweep experiments on the OpenWebText-scale Fock-PARFLM variant confirm that the effective damping γeff is much smaller than the nominal coefficient.
Key Design Properties
Globally conservative: Both \(V_\theta\) and \(V_\phi\) are scalar potentials; the total force \(f = -\nabla(V_\theta + \sum V_\phi)\) is conservative by construction.
Sparse routing: Gumbel-softmax top-k selection keeps pairwise cost at \(O(Tk)\) instead of \(O(T^2)\).
Stage-1.5b gathered \(V_\phi\): Memory-efficient implementation replacing \(O(T^2)\) intermediates with \(O(Tk)\).
The PARFLM is not based on the Transformer architecture. There are no attention layers, no key-value cache, and no feed-forward network towers. The model uses two small scalar-potential MLPs: \(V_\theta\) (single-body, ≈3.4M params) and \(V_\phi\) (pairwise, ≈19K params) whose gradients provide conservative forces. Pairwise interactions use Gumbel-softmax top-k sparse routing at \(O(Tk)\) cost — not \(O(T^2)\) attention.
Because the model carries only a fixed-size state \((h, v, \xi)\) per position — with no KV-cache — its inference memory is \(O(1)\) in sequence length. The figure below illustrates the widening memorization gap between the Transformer's linearly-growing KV-cache and the SPLM's constant-size dynamic state:
Runtime information capacity vs sequence length
Geometric Capabilities of Conservative Architectures
This model is fully attention-free and conservative by construction. Because all forces derive from the gradient of a scalar potential \(V_\theta\), the hidden-state manifold is endowed with a natural damped Riemannian geometry — the layer-dependent Jacobi metric \(\Omega^2_\ell = 2T_\ell \cdot m\) — which is categorically absent from Transformer architectures. This geometry opens the door to capabilities that cannot be replicated in attention-based models:
Capability
Conservative SPLM
Transformer
Riemannian metric on hidden states
Layer-dependent Jacobi metric \(\Omega^2_\ell = 2T_\ell \cdot m\) from \(V_\theta\); confirmed positive at 100% of positions (diagnostic battery Arm 1)
No metric structure
Geodesics between semantic states
Damped geodesic equation with friction term \(-\gamma v\); directional cosine similarity 0.52–0.75 (Arm 2). Geodesics are asymmetric: \(d(A \to B) \neq d(B \to A)\)
Linear interpolation only
Controlled energy dissipation as inference signal
\(\Delta E_{\text{anomaly}}(t) =
\Delta E(t) - \Delta E_{\text{expected}}(t)
Curvature as uncertainty measure
\(\mathcal{K}{\max} = \lambda{\max}(\nabla^2 V_\theta) / 2T_\ell\); well-defined across all layers (Arm 3)
None
These structural properties enable a set of native architectural features that are planned or under investigation (Section 18d and Section 23 of the paper):
Geodesic Analogical Reasoning: Analogy completion via parallel transport of directed geodesic arcs on the semantic manifold, respecting potential barriers that linear embedding arithmetic ignores. The damped geodesic equation yields 3–20% cosine-similarity improvement over undamped (diagnostic battery Arm 2). Because damped geodesics are asymmetric, analogy transport must use directed arcs.
Native Hallucination Detection: Energy dissipation anomalies \(\Delta E_{\text{anomaly}}(t) = |\Delta E(t) - \Delta E_{\text{expected}}(t)|\) and curvature spikes \(\mathcal{K}_{\max}(t)\) provide mechanistically grounded uncertainty signals computable at inference time without additional parameters. The smooth damping-induced energy decay is normal operation; deviation from the expected dissipation curve flags hallucination. For Fock models, the detector needs a per-model baseline that accounts for the known layer-1 exchange transient.
Geodesic Semantic Distance: A replacement for cosine similarity that encodes the model's learned energy landscape, expected to outperform cosine on polysemy and cross-basin semantic cases. The geodesic distance is inherently asymmetric: \(d_{\text{geo}}(A,B) \neq d_{\text{geo}}(B,A)\); a symmetrised variant \([d(A \to B) + d(B \to A)]/2\) is available when symmetry is desired.
Native Chain-of-Thought(via Fock extension): The Fock-PARFLM v2.1 extends this model with register-based native CoT — reasoning steps as Fock register waypoints on damped geodesics, with zero extra token generation. The diagnostic confirms Fock register dynamics are predominantly linear — \(R^2_{\text{full}} \approx 0.83\) (Arm 5), supporting the geodesic-waypoint interpretation.
The conservative constraint imposes a PPL cost relative to attention, but the price buys geometric structure and interpretability that attention-based architectures are structurally incapable of providing.
How to Get Started
python
1# Clone the companion repository for full source code2# git clone https://github.com/dimitarpg13/semsimula-paper.git3# cd semsimula-paper/notebooks/conservative_arch45import torch
6import sys
7sys.path.insert(0,"parf")8sys.path.insert(0,"multixi")9sys.path.insert(0,"energetic_minima")10sys.path.insert(0,"sarf_mass_variant")1112from parf.model_parf_multixi import MultiXiPARFLM, MultiXiPARFConfig
1314config = MultiXiPARFConfig(15 vocab_size=50257,16 d=256,17 n_layers=8,18 v_hidden=1024,19 v_depth=3,20 max_len=1024,21 block_size=512,22 gamma=0.30,23 xi_channels=8,24 xi_alpha_inits="log_spaced",25 v_phi_kind="structural_competitive",26 v_phi_hidden=128,27 top_k=8,28)2930model = MultiXiPARFLM(config)31print(f"Parameters: {sum(p.numel()for p in model.parameters()):,}")3233# Forward pass34x = torch.randint(0,50257,(1,64))35logits, loss = model(x, targets=x)
Available Checkpoint
A trained checkpoint (PPL 12.06, 8k steps) is included in this repository:
File
Description
checkpoint/model.pt
Full model state dict (67 MB)
training_log.jsonl
Per-step training metrics
loss_curve.png
Training/validation loss plot
training_summary.md
Hyperparameters and final metrics
To load the checkpoint:
python
1from huggingface_hub import hf_hub_download
2import torch
34# Download checkpoint5ckpt_path = hf_hub_download(6 repo_id="dimitarpg13/semsimula-parflm-multixi",7 filename="checkpoint/model.pt",8)910# Load into model (after creating model as above)11state = torch.load(ckpt_path, map_location="cpu")12model.load_state_dict(state["model_state_dict"])13model.eval()
Training Details
Training Data
TinyStories -- GPT-2 BPE tokenization. Training cap: 5M tokens.
Training Procedure
Hyperparameter
Value
Optimizer
AdamW
Learning rate
5e-4 (cosine decay)
Warmup steps
400
Weight decay
0.01
Gradient clipping
1.0
Batch size
16
Block size
512
Training steps
8,000
Memory optimisation
Level-2 grad checkpoint + Stage-1.5b gathered \(V_\phi\)
Adding sparse pairwise forces \(V_\phi\) improves PPL from 11.51 to 12.06 at the same step count over the standalone Multi-Xi SPLM — note: the SPLM catches up at 16k steps (11.51 PPL) while PARFLM was only trained to 8k steps, confirming that pairwise token interactions are a necessary complement to the single-body potential. The remaining gap to attention is closed further by the Fock register mechanism.
SPLM Family Overview
This model is part of the Semantic Simulation SPLM family: