LLM Evaluation: Epistemological & Logical Blind Spots in Base Models
1. Model Tested
Model: Nanbeige/Nanbeige4-3B-Base
This is a 3-billion parameter base model developed by the Nanbeige LLM Lab, pre-trained on a comprehensive 23-trillion-token corpus. It lacks supervised fine-tuning (SFT) or Reinforcement Learning from Human Feedback (RLHF) for chat alignment.