NVFP4 quantization of llm-jp/llm-jp-4-8b-instruct
for native serving on NVIDIA Blackwell GPUs with vLLM: RTX PRO 6000 (SM120)
and DGX Spark / GB10 (SM121). It is 8.9 GB (0.52× of BF16), recovers
~99.1% of the BF16 quality score, and improves single-stream decode by
1.76× on RTX PRO 6000 and 2.02× on DGX Spark / GB10 in the
llama-benchy measurements below.
The calibration set is a Japanese-heavy chat/instruction mix matched to the
model's expected deployment distribution. The full recipe, calibration
datasheet, and every evaluation number (including the unflattering ones) are
published here.
Then use any OpenAI-compatible client against /v1/chat/completions — the
bundled plugin registers the Harmony reasoning parser and sets the required
detokenization defaults, so no special request fields are needed.
Quantization recipe
Tool: NVIDIA Model Optimizer (nvidia-modelopt==0.45.0), modelopt-format
export → vLLM's native fast path (FlashInferCutlassNvFp4LinearKernel)
Scale integrity: all FP8-E4M3 weight_scale tensors are byte-scanned for
NaN encodings (0x7F/0xFF) as a standard post-export gate — zero present
(quantization_recipe/scan_weight_scale_nan.py).
Mixed precision (following NVIDIA's own NVFP4 model releases):
First 2 and last 2 decoder layers, embeddings, lm_head, norms: BF16
(Nemotron-style exclusion of the most quantization-sensitive layers)
Calibration: 588 samples (~1.5M tokens):
512 conversations (40% JA multi-turn chat, 20% JA knowledge, 15% JA math,
10% JA-instructed code, 15% EN chat), passed through the model's Harmony
chat template so activation ranges see deployment-realistic token streams.
76 strict-format samples in raw completion form (no chat template):
<answer>-tagged NLI (content from JNLI train — no eval-set overlap)
and short-answer QA, plus JSON-schema outputs. Constrained-output formats
(tags, JSON, fixed labels) are served as raw zero-shot prompts by many
benchmark harnesses and applications, so they are calibrated in that form
too; the slice deliberately uses the same zero-shot instruction scaffold
conventions as Japanese evaluation harnesses (llm-jp-eval style), with
content drawn only from training splits. Datasheet and the exact
calibration file: quantization_recipe/.
Base model revision 098f2b2cf33021eba19a6d3582aa3d071ccc0aff; environment
freeze in quantization_recipe/environment_freeze.txt.
Exact reproduction (base revision and calibration file default to the
pinned/bundled ones):
Evaluation (all measured through served vLLM kernels, not simulation)
Both columns measured on the same harness, same settings, same machine, same
serving stack (RTX PRO 6000 Blackwell, vLLM 0.24.0). llm-jp-eval v2.1.5,
100 samples/dataset, temperature 0, 58 datasets + xlsum supplement.
Metric
BF16
this checkpoint
recovery
llm-jp-eval AVG (14 categories)
0.5409
0.5363
99.1%
ja MT-Bench (3-round mean, fixed local judge)
7.44
7.17
−0.28
JA held-out perplexity (200k tokens)
11.82
11.98
+1.3%
mgsm (full 250, paired)
0.224
0.204
net −5, n.s.
gsm8k (full 1319, paired)
0.224
0.212
net −16 (z=1.6, n.s.)
mawps (full 500, paired)
0.836
0.842
within noise
Category detail (JA, this checkpoint vs BF16): NLI 0.800/0.810,
QA 0.479/0.506, RC 0.710/0.710, CR 0.870/0.867, HE 0.383/0.385,
EL 0.600/0.614, FA 0.314/0.326, MR 0.341/0.368, MT 0.842/0.842,
IF 0.490/0.450, BBH 0.605/0.585, LM 0.780/0.800; EN AVG 0.444/0.454.
Math is evaluated on full test sets with paired per-item flip analysis
(McNemar-style) because n=100 binomial noise cannot resolve ~1% relative
changes; "n.s." = not statistically significant.
MT-Bench note: judged with a fixed local judge (Gemma-4-31B-IT-NVFP4, greedy),
all rounds of both models in one batch session, generation max_tokens=2048.
Round-to-round generation spread is ±0.04–0.12. Scores are relative
comparisons only — not comparable to LLM-jp's published GPT-judge numbers.
Throughput: BF16 vs NVFP4
Measured with llama-benchy, vLLM 0.24.0, 2,048 prompt tokens, 256 generated
tokens, and three measured runs.
Hardware
SM
concurrency
BF16 decode tok/s
NVFP4 decode tok/s
speedup
TTFT (e2e)
RTX PRO 6000
SM120
1
84.4
148.1
1.76×
155→91 ms
RTX PRO 6000
SM120
4
297
368
1.24×
438→238 ms
RTX PRO 6000
SM120
8
450
833
1.85×
847→439 ms
RTX PRO 6000
SM120
16
801
1170
1.46×
1627→808 ms
DGX Spark / GB10
SM121
1
13.80
27.84
2.02×
648→348 ms
DGX Spark / GB10
SM121
4
56.49
106.13
1.88×
1775→1084 ms
DGX Spark / GB10
SM121
8
97.36
173.16
1.78×
2900→1597 ms
DGX Spark / GB10
SM121
16
153.11
253.57
1.66×
4763→2638 ms
On RTX PRO 6000, c=1 decode is stable to <0.1% across runs; the c=4–16
points carry ±10–15% run-to-run spread. KV-cache capacity at
--gpu-memory-utilization 0.9 grows ~15% thanks to the smaller weights.
The DGX Spark / GB10 rows were measured locally on 2026-07-26 against BF16
llm-jp/llm-jp-4-8b-instruct revision
098f2b2cf33021eba19a6d3582aa3d071ccc0aff and this NVFP4 checkpoint revision
748431048bb5b7114424ae6c0a17bdd2172cc9cf.
On DGX Spark / GB10 (SM121), a Japanese ELYZA-100 streaming decode benchmark
measured 14.42 tok/s for BF16, 29.73 tok/s for this NVFP4 checkpoint, and
49.04 tok/s for NVFP4 + DFlash. That is a 3.40× mean decode throughput
increase from BF16 to NVFP4 + DFlash, with DFlash adding 1.64× over the NVFP4
verifier baseline in the paired run.
Serving notes
Requires vLLM ≥ 0.24 (modelopt mixed-precision NVFP4 is mainline in
stable ≥ 0.24; the load path selects quantization=modelopt_mixed with
FlashInferCutlassNvFp4LinearKernel). vLLM 0.24.0 is the recommended and
benchmark-validated version. 0.25.0 loads and scores identically in our
testing (same kernels, +2% decode), but we observed one unreproduced
cudaErrorIllegalAddress engine crash under high-concurrency load on it and
haven't cleared it for production yet. Validated end-to-end on SM120
(RTX PRO 6000) and SM121 (DGX Spark / GB10).
The base model emits its responses in OpenAI-Harmony structure
(<|channel|>analysis/final<|message|>…), which is why the bundled plugin is
needed: it registers the llmjp4 reasoning parser (splitting reasoning from
content), and when that parser is selected it defaults
skip_special_tokens to false on chat completions so the Harmony structure
survives detokenization. An explicit request value always wins, and other
models/parsers on a server with the plugin installed are unaffected. For
setups where the plugin cannot be installed server-side, a transparent proxy
with the same effect is included (UPSTREAM_BASE_URL=http://localhost:8000 PROXY_PORT=8080 python3 vllm_plugin/harmony_proxy.py).
The plugin's parser also handles a SentencePiece boundary-space artifact in
vLLM's incremental detokenizer that otherwise breaks non-streaming channel
extraction.
Known limitations (measured, not hypothetical)
Small math regression: paired per-item analysis on the full gsm8k test
set shows a net −16/1319 items vs BF16 (z=1.6, not statistically
significant). mgsm and mawps show no significant change.
QA softness: JA QA category 94.6% of BF16, EN QA ~91% (drop/mmlu-style
extractive QA; n=100/dataset).
Base-model quirk (not quantization-related): the model may identify itself
as "ChatGPT" when asked who it is.
License & attribution
Apache-2.0, same as the base model. Base model by
LLM-jp. Calibration data: llm-jp/oasst2-33k-ja
(Apache-2.0), wikimedia/wikipedia 20231101.ja (CC-BY-SA-4.0),
Kendamarron/magpie-japanese-math-instruction-17k-qwen2.5-bakeneko-32b-instruct (Apache-2.0),
ronantakizawa/python-code-instructions-japanese (MIT),
HuggingFaceH4/ultrachat_200k (MIT), JGLUE/JNLI train split via
zenless-lab/jnli (CC-BY-SA-4.0). The CC-BY-NC xlsum_ja dataset was used ONLY
for evaluation scoring, never for calibration.
Base models are trained with pre-training and mid-training only.
Post-trained models are aligned using supervised fine-tuning (SFT) and direct preference optimization (DPO), without reinforcement learning.
[!NOTE]
While the thinking variants are trained with both SFT and DPO, this instruct model is trained using SFT only, without DPO.
For practical usage examples and detailed instructions on how to use the models, please also refer to our cookbook.
To support the continued development of LLM-jp, we would greatly appreciate it if you could share how you utilize LLM-jp outcomes via the survey form.
Usage
Please refer to our cookbook for practical usage examples and detailed instructions on how to use the models.
Model Details
Model type: Transformer-based Language Model
Architectures:
Dense model:
Params
Layers
Hidden size
Heads
Context length
Embedding parameters
Non-embedding parameters
Total parameters
8B
32
4,096
32
65,536
805,306,368
7,784,894,464
8,590,200,832
MoE model:
Params
Layers
Hidden size
Heads
Routed Experts
Activated Experts
Context length
Embedding parameters
Non-embedding parameters
Activated parameters
Total parameters
32B-A3B
32
2,560
40
128
8
65,536
503,316,480
31,635,712,512
3,827,476,992
32,139,028,992
Tokenizer
The tokenizer of this model is based on huggingface/tokenizers Unigram byte-fallback model.
The vocabulary entries were converted from llm-jp-tokenizer v4.0.
Please refer to README.md of llm-jp-tokenizer for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our vocabulary).
[!NOTE]
The chat template of this model is designed to be compatible with the OpenAI Harmony response format.
However, the tokenizer differs from the one assumed by the openai-harmony library, and therefore direct tokenization with openai-harmony is not supported.
For correct behavior, please use the tokenizer provided with this model. For detailed usage, please refer to our cookbook.
Training
Pre-training
This model is trained through a multi-stage pipeline consisting of pre-training and mid-training phases, using a total of 11.7T tokens.
pretraining_overview
The corpora used for pre-training and mid-training are publicly available at the following links:
We evaluated the model on a variety of tasks using an LLM-as-a-Judge framework. The descriptions of each task are as follows.
MT-Bench (JA/EN): A benchmark for measuring multi-turn conversational task-solving ability.
AnswerCarefully: A benchmark for evaluating safety in Japanese. We used 336 questions from the v2.0 test set.
llm-jp-instructions: A set of human-created single-turn question–answer pairs. We used 400 questions from the test set.
We evaluated the models using gpt-5.4-2026-03-05.
[!NOTE]
Note: In earlier evaluations of the llm-jp-3 series, we used gpt-4o-2024-08-06. The newer evaluator gpt-5.4-2026-03-05 provides a stricter and more reliable assessment, which results in lower scores on benchmarks such as MT-Bench compared to those reported for the llm-jp-3 series.
The scores represent the average values obtained from three rounds of inference and evaluation.
For more details, please refer to the codes.
The models released here are in the early stages of our research and development and have not been tuned to ensure outputs align with human intent and safety considerations.