Two modified versions of Qwen3.5-27B produced by
RYS layer duplication — no training, no weight changes, just routing hidden states through a specific circuit twice.
Scores from an internal sweep benchmark run during circuit search. Sample sizes are small — treat these as directional indicators, not definitive benchmarks.
Model
Math
EQ
Reasoning
Logic
Base (64 layers)
0.375
11.5
0.000
0.00
rys_30-33 (68 layers)
0.438
29.5
0.353
1.00
rys_34-37 (68 layers)
0.375
39.4
0.000
0.00
Math: Ng's partial-credit scoring on a small GSM8K sample
Reasoning: fraction correct across causal, date, logic, navigation, and GSM8K probes
Logic: fraction correct on logical deduction probes only
rys_30-33 shows the best combined improvement across reasoning categories. rys_34-37 shows the highest EQ score but no reasoning improvement over baseline.
Benchmarks (based on BFCLv4)
Non-Live Tests
Task
Qwen3.5-27B-RYS-30-34 (Δ vs Best)
Qwen3.5-27B-FC (Baseline)
Claude Opus 4.5 (FC)
Claude Sonnet 4.5 (FC)
GLM 4.6 (FC)
Grok-4 (FC)
GPT-5.2 (FC)
irrelevance
86.67% (-1.25%)
87.50%
85.83%
87.92%
85.42%
77.50%
80.00%
multiple
96.50%
96.50%
95.50%
95.50%
95.00%
92.50%
88.00%
parallel
95.00%
93.00%
93.50%
94.50%
91.50%
88.50%
89.00%
parallel_multiple
91.50% (-0.50%)
76.00%
88.50%
92.00%
89.50%
87.00%
77.50%
simple_java
62.00% (-3.00%)
65.00%
60.00%
62.00%
64.00%
62.00%
62.00%
simple_javascript
72.00% (-2.00%)
66.00%
74.00%
58.00%
64.00%
66.00%
64.00%
simple_python
95.25% (-2.50%)
95.00%
96.50%
97.75%
94.75%
92.50%
92.75%
Live Tests
Task
Qwen3.5-27B-RYS-30-34 (Δ vs Best)
Qwen3.5-27B-FC (Baseline)
Claude Opus 4.5 (FC)
Claude Sonnet 4.5 (FC)
GLM 4.6 (FC)
Grok-4 (FC)
GPT-5.2 (FC)
live_irrelevance
82.24% (-3.05%)
80.88%
83.60%
85.29%
84.50%
73.30%
78.85%
live_multiple
79.68% (-1.14%)
80.82%
78.16%
78.92%
78.92%
73.88%
70.37%
live_parallel
81.25% (-6.25%)
87.50%
87.50%
87.50%
81.25%
75.00%
68.75%
live_parallel_multiple
75.00% (-8.33%)
79.17%
75.00%
83.33%
75.00%
79.17%
58.33%
live_relevance
81.25% (-6.25%)
68.75%
62.50%
68.75%
75.00%
87.50%
75.00%
live_simple
84.50% (-5.03%)
87.60%
86.43%
89.53%
89.53%
82.17%
71.71%
Multi-Turn Tests
Task
Qwen3.5-27B-RYS-30-34 (Δ vs Best)
Qwen3.5-27B-FC (Baseline)
Claude Opus 4.5 (FC)
Claude Sonnet 4.5 (FC)
GLM 4.6 (FC)
Grok-4 (FC)
GPT-5.2 (FC)
multi_turn_base
74.50% (-6.50%)
70.50%
81.00%
69.00%
74.50%
44.00%
36.50%
multi_turn_long_context
67.50% (-3.00%)
59.00%
70.50%
59.00%
66.50%
44.00%
30.50%
Memory Tests (Agentic)
Task
Qwen3.5-27B-RYS-30-34 (Δ vs Best)
Qwen3.5-27B-FC (Baseline)
Claude Opus 4.5 (FC)
Claude Sonnet 4.5 (FC)
GLM 4.6 (FC)
Grok-4 (FC)
GPT-5.2 (FC)
memory_kv
45.81% (-25.16%)
N/A
70.97%
54.19%
43.87%
57.42%
33.55%
memory_rec_sum
70.97% (-12.26%)
N/A
77.42%
83.23%
67.10%
51.61%
60.65%
memory_vector
63.23% (-9.67%)
N/A
72.90%
57.42%
56.13%
58.71%
43.23%
RYS vs Baseline Comparison (All Tests)
Task
RYS
Baseline
Δ (RYS - Baseline)
irrelevance
86.67%
87.50%
-0.83%
multiple
96.50%
96.50%
0.00%
parallel
95.00%
93.00%
+2.00% ✅
parallel_multiple
91.50%
76.00%
+15.50% ✅
simple_java
62.00%
65.00%
-3.00%
simple_javascript
72.00%
66.00%
+6.00% ✅
simple_python
95.25%
95.00%
+0.25%
live_irrelevance
82.24%
80.88%
+1.36% ✅
live_multiple
79.68%
80.82%
-1.14%
live_parallel
81.25%
87.50%
-6.25%
live_parallel_multiple
75.00%
79.17%
-4.17%
live_relevance
81.25%
68.75%
+12.50% ✅
live_simple
84.50%
87.60%
-3.10%
multi_turn_base
74.50%
70.50%
+4.00% ✅
multi_turn_long_context
67.50%
59.00%
+8.50% ✅
memory_kv
45.81%
N/A
✅
memory_rec_sum
70.97%
N/A
✅
memory_vector
63.23%
N/A
✅
What is RYS?
Transformers self-organise during training into functional circuits — contiguous blocks of layers that act together. The RYS technique duplicates a specific block in the forward pass using the same weights, with no extra copies on disk beyond the GGUF file overhead:
The model weights alone are ~21 GiB (Q4_K_XL quantization, 68 layers). A single A100 80GB or H100 runs this comfortably. Consumer GPU setups depend on your llama.cpp version's tensor split support.