Manifest 4B · v0.6.0
The most capable Manifest model available today — deep reasoning for the hardest, multi-constraint supply-chain problems.
Manifest is MetaFloor's suite of supply-chain expert models — purpose-built specialists in procurement,
demand planning, warehouse operations, supplier relationship management, risk & resilience,
transportation, inventory, and order fulfilment.
- Family: Manifest · This model: Manifest 4B (4-billion-parameter base)
- Tier: advanced — the most capable model available today
Preferred 95.1% of the time over the base model
On 134 held-out expert questions, an independent LLM judge panel picked this model's answer over the
untuned base model's answer 95.1% of the time (95% CI 91.0–98.5%). With both models given the same
answer format, it is still preferred 97.0% of the time — the gain is real domain knowledge, not just presentation.
What's new in v0.6
- Retrained on MetaFloor's expanded ~32k-example supply-chain dataset (up from ~12.5k in v0.5).
- The evaluation benchmark grew to 134 held-out questions (from 116) — so v0.6 headline figures are measured on a larger, harder set than the v0.5 cards.
- A new Manifest 9B tier is joining the family (coming soon); the 35B orchestrator is retired from this release.
- This is a new versioned repo — the previous v0.5 release (
metafloor-ai/manifest-orchestrator-4b, unchanged) keeps its own history and download stats.
v0.5 → v0.6, same 134-question benchmark (preferred over the untuned base):
| Release | Preferred over base | vs base + matched format |
|---|
| v0.5.0 | 88.8% | 90.9% |
| v0.6.0 | 95.1% | 97.0% |
(The baseline throughout is the untuned Qwen/Qwen3.5-4B — every win-rate above is measured against it.)
Best for
The heavy end: complex, multi-constraint scenarios where the answer has to be right and reasoned.
- Multi-echelon strategy, network design, and cost-vs-service trade-offs
- Executive-level advisory over messy, real-world operating constraints
- Nuanced supplier, risk, and resilience decisions
- Any task where answer quality matters more than latency or cost
(Need speed or volume instead? See Manifest 2B and Manifest 0.8B below.)
See the difference
Same question. Base model vs Manifest 4B.
Ask (COO of a $5.8B enterprise, 176 nodes, mostly manual): "Explain the core concept of MEIO and why it's relevant given our manual, single-echelon approach."
Base model → "Thinking Process: 1. Analyze the Request: Role: COO of a global enterprise. Company Size: $5,812M revenue. Network: 176 network nodes…" — narrates its own reasoning instead of answering.
Manifest 4B → "MEIO is the optimization of inventory across multiple echelons of the supply chain at once, rather than optimizing each echelon separately…" — answers directly, like a practitioner, then ties it to the operation.
The Manifest family
Two kinds of models:
🧭 Orchestrators — general-purpose, handle any supply-chain area
| Model | Size | Preferred over base | Status |
|---|
| Manifest 0.8B | 0.8B | 79.9% | ✅ available |
| Manifest 2B | 2B | 85.1% | ✅ available |
| Manifest 4B | 4B | 95.1% | ✅ available |
| Manifest 9B | 9B | — | 🔜 coming soon |
🎯 Domain-experts — specialized for a single area
| Model | Preferred over base | Status |
|---|
| Manifest Specialist · Risk & Resilience | 100% (20 / 0) | ✅ available |
| Manifest Specialist · Inventory Optimization | 80.0% | ✅ available |
| Manifest Specialist · Demand Planning | 82.5% | ✅ available |
Orchestrators are scored on the general supply-chain benchmark; domain-experts on their focused domain benchmark (20 items each — treat those exact figures as directional).
Not just a model — a full stack
Manifest is backed by everything needed to build it and trust it:
- A purpose-built dataset — thousands of supply-chain instruction–response pairs spanning 8 sub-domains and every company scale, generated by a seed-driven operator-as-teacher pipeline.
- A reproducible training pipeline — documented LoRA fine-tuning.
- An independent benchmark — 134 held-out expert questions, scored blind by a panel of LLM judges.
We built the model, the data, and the evaluation.
How to use
Manifest 4B is a LoRA adapter (~101 MB), applied on top of its base model at load time.
1from peft import PeftModel
2from transformers import AutoModelForCausalLM, AutoTokenizer
3
4base = "Qwen/Qwen3.5-4B" # base model — see "Built on" below
5tok = AutoTokenizer.from_pretrained(base)
6model = AutoModelForCausalLM.from_pretrained(base, device_map="auto")
7model = PeftModel.from_pretrained(model, "metafloor-ai/manifest-orchestrator-4b-v0.6.0")
8
9SYSTEM = "You are a senior supply chain expert. Answer correctly and concisely."
10user = (
11 "I'm an inventory planner at a ~$8M small business: ~11k active SKUs, 4 suppliers, "
12 "2 network nodes, ~164-day avg lead time. How should I set safety stock as I move off spreadsheets?"
13)
14msgs = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": user}]
15inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_dict=True, return_tensors="pt")
16out = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
17print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Prompt tip: Manifest is trained to condition on the scenario — include the asker's role and operating
scale (revenue, SKUs, suppliers, nodes, lead time) in the message for the sharpest, most tailored answers.
How it was measured
134 held-out expert questions across 8 supply-chain areas. Each question is answered by Manifest and by
the base model (given the same answer format); an independent two-model LLM judge panel then picks the
better answer. Manifest 4B was preferred
97.0% of the time (95% CI 94.0–99.3%; 129 wins / 3 losses / 2 ties over 134). Benchmark:
supply-chain-eval.
Training details
| |
|---|
| Method | LoRA (PEFT 0.20.0), rank 16 / alpha 16 / dropout 0.05 |
| Target modules | all attention + MLP projections |
| Trainable params | 21,233,664 (~0.87% of the 2.44B base) |
| Epochs | 3 |
| Training examples | ~32,000 |
| Final loss | 1.09 (from 2.28) |
Training data: supply-chain instruction–response pairs from the seed-driven
operator-as-teacher
pipeline — a deterministic engine emits a unique seed per example (area, sub-area, persona, question
type, realistic numeric scenario) and a strong teacher model writes the matching answer.
The training data is drawn from MetaFloor's proprietary ~32k-example supply-chain dataset, which is not open-sourced — only the held-out evaluation benchmark (
supply-chain-eval) is public.
Intended use & limitations
- Intended use: high-quality decision-support and drafting for supply-chain professionals.
- Out of scope: not legally binding, contractual, or safety-critical guidance; no access to your live
systems or real-time data. Verify outputs before acting on them.
- Limitations: English-only; trained on synthetic (model-authored) data; standard LLM risks
(hallucination, outdated facts) apply.
License
Manifest models and the
supply-chain-eval benchmark are released under
CC-BY-NC-4.0 — free for research and non-commercial use, with attribution.
Commercial use requires a license from MetaFloor — get in touch at
metafloor.ai.
Built on
Manifest 4B is a LoRA adapter over Qwen/Qwen3.5-4B (used under its own license); the base model is required to load the adapter.
Citation
1@misc{metafloor_manifest_4b,
2 title = {Manifest 4B: a supply-chain expert model (MetaFloor Manifest suite)},
3 author = {MetaFloor AI},
4 year = {2026},
5 howpublished = {\url{https://huggingface.co/metafloor-ai/manifest-orchestrator-4b-v0.6.0}}
6}