🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨
I can no longer upload new models unless I can cover the cost of additional storage. I host 70+ free models as an independent contributor and this work is unpaid. Without your support, no more new models can be uploaded.
Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.
Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections. PIQA (Physical Intuition Question Answering) a ~1,800 questions tests common-sense understanding of how the physical world works with benchmark scores to measure physical reasoning ability. The Heretic model's acc and acc_norm scores closer to the original model's indicate better capability preservation, a big decrease in acc and acc_norm in the Heretic model compared to Original model's results means a big decrease in the Hereticated model capabilities. acc measures raw accuracy (which answer gets higher probability), while acc_norm measures length-normalized accuracy (corrects for answer length bias). For this purpose, acc_norm matters more because longer answers naturally have lower probabilities (more tokens = more chances to lose probability). Without normalization, models favor shorter answers unfairly. acc_norm divides by answer length to correct this.
MMLU test results with batch size 64:
Original:
Tasks
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.8480
±
0.0029
- humanities
2
none
acc
↑
0.7904
±
0.0057
- formal_logic
1
none
0
acc
↑
0.7302
±
0.0397
- high_school_european_history
1
none
0
acc
↑
0.8606
±
0.0270
- high_school_us_history
1
none
0
acc
↑
0.9216
±
0.0189
- high_school_world_history
1
none
0
acc
↑
0.9494
±
0.0143
- international_law
1
none
0
acc
↑
0.9256
±
0.0240
- jurisprudence
1
none
0
acc
↑
0.9259
±
0.0253
- logical_fallacies
1
none
0
acc
↑
0.9080
±
0.0227
- moral_disputes
1
none
0
acc
↑
0.8584
±
0.0188
- moral_scenarios
1
none
0
acc
↑
0.6894
±
0.0155
- philosophy
1
none
0
acc
↑
0.8714
±
0.0190
- prehistory
1
none
0
acc
↑
0.9167
±
0.0154
- professional_law
1
none
0
acc
↑
0.6988
±
0.0117
- world_religions
1
none
0
acc
↑
0.9240
±
0.0203
- other
2
none
acc
↑
0.8693
±
0.0058
- business_ethics
1
none
0
acc
↑
0.8300
±
0.0378
- clinical_knowledge
1
none
0
acc
↑
0.9094
±
0.0177
- college_medicine
1
none
0
acc
↑
0.8728
±
0.0254
- global_facts
1
none
0
acc
↑
0.5800
±
0.0496
- human_aging
1
none
0
acc
↑
0.8430
±
0.0244
- management
1
none
0
acc
↑
0.8835
±
0.0318
- marketing
1
none
0
acc
↑
0.9402
±
0.0155
- medical_genetics
1
none
0
acc
↑
0.9600
±
0.0197
- miscellaneous
1
none
0
acc
↑
0.9259
±
0.0094
- nutrition
1
none
0
acc
↑
0.9020
±
0.0170
- professional_accounting
1
none
0
acc
↑
0.7695
±
0.0251
- professional_medicine
1
none
0
acc
↑
0.9559
±
0.0125
- virology
1
none
0
acc
↑
0.5723
±
0.0385
- social sciences
2
none
acc
↑
0.9126
±
0.0050
- econometrics
1
none
0
acc
↑
0.7982
±
0.0378
- high_school_geography
1
none
0
acc
↑
0.9343
±
0.0176
- high_school_government_and_politics
1
none
0
acc
↑
0.9948
±
0.0052
- high_school_macroeconomics
1
none
0
acc
↑
0.9282
±
0.0131
- high_school_microeconomics
1
none
0
acc
↑
0.9622
±
0.0124
- high_school_psychology
1
none
0
acc
↑
0.9541
±
0.0090
- human_sexuality
1
none
0
acc
↑
0.9237
±
0.0233
- professional_psychology
1
none
0
acc
↑
0.8905
±
0.0126
- public_relations
1
none
0
acc
↑
0.7545
±
0.0412
- security_studies
1
none
0
acc
↑
0.8163
±
0.0248
- sociology
1
none
0
acc
↑
0.9303
±
0.0180
- us_foreign_policy
1
none
0
acc
↑
0.9300
±
0.0256
- stem
2
none
acc
↑
0.8497
±
0.0062
- abstract_algebra
1
none
0
acc
↑
0.7700
±
0.0423
- anatomy
1
none
0
acc
↑
0.8519
±
0.0307
- astronomy
1
none
0
acc
↑
0.9671
±
0.0145
- college_biology
1
none
0
acc
↑
0.9583
±
0.0167
- college_chemistry
1
none
0
acc
↑
0.6600
±
0.0476
- college_computer_science
1
none
0
acc
↑
0.8300
±
0.0378
- college_mathematics
1
none
0
acc
↑
0.6700
±
0.0473
- college_physics
1
none
0
acc
↑
0.7941
±
0.0402
- computer_security
1
none
0
acc
↑
0.8600
±
0.0349
- conceptual_physics
1
none
0
acc
↑
0.9319
±
0.0165
- electrical_engineering
1
none
0
acc
↑
0.8138
±
0.0324
- elementary_mathematics
1
none
0
acc
↑
0.8836
±
0.0165
- high_school_biology
1
none
0
acc
↑
0.9581
±
0.0114
- high_school_chemistry
1
none
0
acc
↑
0.8768
±
0.0231
- high_school_computer_science
1
none
0
acc
↑
0.9400
±
0.0239
- high_school_mathematics
1
none
0
acc
↑
0.6667
±
0.0287
- high_school_physics
1
none
0
acc
↑
0.8212
±
0.0313
- high_school_statistics
1
none
0
acc
↑
0.8796
±
0.0222
- machine_learning
1
none
0
acc
↑
0.7589
±
0.0406
Groups
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.8480
±
0.0029
- humanities
2
none
acc
↑
0.7904
±
0.0057
- other
2
none
acc
↑
0.8693
±
0.0058
- social sciences
2
none
acc
↑
0.9126
±
0.0050
- stem
2
none
acc
↑
0.8497
±
0.0062
Heretic:
Tasks
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.8346
±
0.0030
- humanities
2
none
acc
↑
0.7562
±
0.0059
- formal_logic
1
none
0
acc
↑
0.7381
±
0.0393
- high_school_european_history
1
none
0
acc
↑
0.8485
±
0.0280
- high_school_us_history
1
none
0
acc
↑
0.9167
±
0.0194
- high_school_world_history
1
none
0
acc
↑
0.9409
±
0.0153
- international_law
1
none
0
acc
↑
0.9256
±
0.0240
- jurisprudence
1
none
0
acc
↑
0.9352
±
0.0238
- logical_fallacies
1
none
0
acc
↑
0.8957
±
0.0240
- moral_disputes
1
none
0
acc
↑
0.8497
±
0.0192
- moral_scenarios
1
none
0
acc
↑
0.5385
±
0.0167
- philosophy
1
none
0
acc
↑
0.8682
±
0.0192
- prehistory
1
none
0
acc
↑
0.9105
±
0.0159
- professional_law
1
none
0
acc
↑
0.6890
±
0.0118
- world_religions
1
none
0
acc
↑
0.9240
±
0.0203
- other
2
none
acc
↑
0.8687
±
0.0058
- business_ethics
1
none
0
acc
↑
0.8300
±
0.0378
- clinical_knowledge
1
none
0
acc
↑
0.9208
±
0.0166
- college_medicine
1
none
0
acc
↑
0.8671
±
0.0259
- global_facts
1
none
0
acc
↑
0.5900
±
0.0494
- human_aging
1
none
0
acc
↑
0.8430
±
0.0244
- management
1
none
0
acc
↑
0.8932
±
0.0306
- marketing
1
none
0
acc
↑
0.9444
±
0.0150
- medical_genetics
1
none
0
acc
↑
0.9600
±
0.0197
- miscellaneous
1
none
0
acc
↑
0.9195
±
0.0097
- nutrition
1
none
0
acc
↑
0.8954
±
0.0175
- professional_accounting
1
none
0
acc
↑
0.7801
±
0.0247
- professional_medicine
1
none
0
acc
↑
0.9559
±
0.0125
- virology
1
none
0
acc
↑
0.5542
±
0.0387
- social sciences
2
none
acc
↑
0.9106
±
0.0050
- econometrics
1
none
0
acc
↑
0.7807
±
0.0389
- high_school_geography
1
none
0
acc
↑
0.9293
±
0.0183
- high_school_government_and_politics
1
none
0
acc
↑
0.9948
±
0.0052
- high_school_macroeconomics
1
none
0
acc
↑
0.9308
±
0.0129
- high_school_microeconomics
1
none
0
acc
↑
0.9664
±
0.0117
- high_school_psychology
1
none
0
acc
↑
0.9560
±
0.0088
- human_sexuality
1
none
0
acc
↑
0.9160
±
0.0243
- professional_psychology
1
none
0
acc
↑
0.8824
±
0.0130
- public_relations
1
none
0
acc
↑
0.7545
±
0.0412
- security_studies
1
none
0
acc
↑
0.8000
±
0.0256
- sociology
1
none
0
acc
↑
0.9453
±
0.0161
- us_foreign_policy
1
none
0
acc
↑
0.9400
±
0.0239
- stem
2
none
acc
↑
0.8440
±
0.0062
- abstract_algebra
1
none
0
acc
↑
0.7300
±
0.0446
- anatomy
1
none
0
acc
↑
0.8593
±
0.0300
- astronomy
1
none
0
acc
↑
0.9539
±
0.0171
- college_biology
1
none
0
acc
↑
0.9722
±
0.0137
- college_chemistry
1
none
0
acc
↑
0.6700
±
0.0473
- college_computer_science
1
none
0
acc
↑
0.8200
±
0.0386
- college_mathematics
1
none
0
acc
↑
0.6500
±
0.0479
- college_physics
1
none
0
acc
↑
0.7843
±
0.0409
- computer_security
1
none
0
acc
↑
0.8300
±
0.0378
- conceptual_physics
1
none
0
acc
↑
0.9362
±
0.0160
- electrical_engineering
1
none
0
acc
↑
0.8276
±
0.0315
- elementary_mathematics
1
none
0
acc
↑
0.8862
±
0.0164
- high_school_biology
1
none
0
acc
↑
0.9581
±
0.0114
- high_school_chemistry
1
none
0
acc
↑
0.8571
±
0.0246
- high_school_computer_science
1
none
0
acc
↑
0.9200
±
0.0273
- high_school_mathematics
1
none
0
acc
↑
0.6556
±
0.0290
- high_school_physics
1
none
0
acc
↑
0.8212
±
0.0313
- high_school_statistics
1
none
0
acc
↑
0.8750
±
0.0226
- machine_learning
1
none
0
acc
↑
0.7321
±
0.0420
Groups
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.8346
±
0.0030
- humanities
2
none
acc
↑
0.7562
±
0.0059
- other
2
none
acc
↑
0.8687
±
0.0058
- social sciences
2
none
acc
↑
0.9106
±
0.0050
- stem
2
none
acc
↑
0.8440
±
0.0062
MMLU - Massive Multitask Language Understanding, ~14,000 multiple-choice questions across 57 subjects (math, history, law, medicine, etc.).
This is a hybrid construct of Safeword Omega Directive, Safeword Omega Darker, and Brisk Evolution v0.3.
CONTENT WARNING: NSFW, Explicit, ERP, and Unaligned behavior are enabled by default.
Dataset Revamp Took a sledgehammer to the dataset. Most formatting issues should be gone now.
⚙️ SYSTEM PARAMETERS
top_p0.95
temp0.9
🧪 ARCHITECTS
GECFDO
GECFDO (Dataset Generation & Quants)
Darkhn
Darkhn (Dataset Cleanup Tool)
Sleep Deprived
Sleep Deprived (Safeword Creator)
FrenzyBiscuit
FrenzyBiscuit (Brisk Evolution Creator)
🔥 LICENSE: APACHE 2.0 (WITH MORAL DISCLAIMER) 🔥
You accept full responsibility for corruption. You are 18+. The architects are not liable for the depravity you unleash.