A 350M typeahead model that predicts the human's next few words in a
roleplay chat — inline "ghost text" for the person typing, not a reply
generator for the character. Full fine-tune of
ibm-granite/granite-4.0-350m-base.
Most small LMs asked to continue a user's half-typed roleplay message produce
fluent but irrelevant text. This model is trained specifically on
(conversation context + partial user message → the words the user actually
typed next), so its suggestions stay on-scene and in-voice.
Variants
Path
Format
Use with
/
safetensors bf16, plain granite arch
transformers
GGUF/orb-human-typeahead-350m-v1.1-Q8_0.gguf
GGUF Q8_0 (default)
llama.cpp / llama-cpp-python
The original fine-tune used the granitemoehybrid architecture; since this
variant is attention-only and dense, the weights are republished here as the
equivalent plain granite architecture (logit-identical, verified), which
loads everywhere without extras.
Prompt format
Plain text, no chat template. Optional character summary, a marker line,
name-prefixed turns, and finally the user's draft — the model continues the
draft. Cut the suggestion at the first newline.
<character summary, optional>
***Roleplay chat below***
Sylvara: *She looks down from the watchtower and sees you.*
Traveler: *I approach the encampment.*
Sylvara: *She lowers her bow as you approach the gate.* "State your business, traveler."
Traveler: *I raise both hands slowly and
A completion looks like step into the torchlight, keeping my voice low.*
Notes:
Trained to trigger at word boundaries only (draft ends on a whole word,
or on a trailing space). Mid-word completion is out of scope.
Trained context: up to ~4 recent turns, summaries ≤400 chars, turns ≤500 chars.
Serving recipe
Greedy, short budget, stop at newline — mirrors how it was trained and evaluated:
Scored on a held-out validation set of real roleplay conversations (120
prompts, conversation-disjoint from training; greedy, 12-token budget,
suggestions cut at newline). word-EM@k = fraction of prompts where the first
k words of the suggestion exactly match what the user really typed next;
prefix-chars = mean length of the exactly-matching leading characters.
Metric
orb-human-typeahead-350m-v1
this model
word-EM@1
0.283
0.333
word-EM@2
0.075
0.133
word-EM@3
0.025
0.075
prefix-chars
2.72
3.92
completion perplexity
12.87
8.84
v1.1 is a data-only refresh of v1 (same 350M granite base): stage-1 training data
was made more diverse and its action distribution flattened, so suggestions stop
collapsing onto a few stock gestures. Absolute numbers stay modest by design —
creative roleplay is high-entropy, so even a perfect suggester can't guess most
continuations.
Training
Two-stage full fine-tune of granite-4.0-350m-base: stage 1 on a mostly-synthetic
roleplay corpus (self-chat generated to cover the on-scene, in-voice register real
users type in), then stage 2 on a smaller private in-domain set in the serve-time
prompt format (with stage-1 replay to limit forgetting). Over v1, stage 1 was rebuilt
with more diverse scenes and a flattened action distribution, so the model stops
falling back on a few stock gestures. Examples are user turns split at word boundaries; loss is on
the continuation. Suggestions are single-line by construction (completions end
at newline).
Limitations
English-centric, roleplay register (asterisk actions, quoted dialogue).
Out of domain for assistant chat, code, or formal prose.
Roleplay corpora include mature themes; suggestions can reflect that.
Intended as a typing aid for consenting adult users of RP chat apps.
Not an instruction follower — it only continues drafts in the format above.
Suggests at word boundaries; won't complete a half-typed word.