Views
No views yet
fable-traces is tuned for short, conversational replies and runs comfortably on a
single mid-range GPU.1from transformers import AutoModelForCausalLM, AutoTokenizer
2import torch
3
4repo = "AliesTaha/fable-traces"
5tok = AutoTokenizer.from_pretrained(repo)
6model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
7
8messages = [{"role": "user", "content": "Tell me something interesting."}]
9ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
10out = model.generate(ids, max_new_tokens=100, do_sample=False)
11print(tok.decode(out[0, ids.shape[1]:], skip_special_tokens=True))vllm serve AliesTaha/fable-traces| Base model | Qwen3-4B-Instruct-2507 |
| Parameters | ~4B |
| Precision | bfloat16 (safetensors) |
| Prompt format | ChatML — use the tokenizer's chat template |
| Context length | inherits the base model |