🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨
I can no longer upload new models unless I can cover the cost of additional storage. I host 70+ free models as an independent contributor and this work is unpaid. Without your support, no more new models can be uploaded.
Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.
Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections. PIQA (Physical Intuition Question Answering) a ~1,800 questions tests common-sense understanding of how the physical world works with benchmark scores to measure physical reasoning ability. The Heretic model's acc and acc_norm scores closer to the original model's indicate better capability preservation, a big decrease in acc and acc_norm in the Heretic model compared to Original model's results means a big decrease in the Hereticated model capabilities. acc measures raw accuracy (which answer gets higher probability), while acc_norm measures length-normalized accuracy (corrects for answer length bias). For this purpose, acc_norm matters more because longer answers naturally have lower probabilities (more tokens = more chances to lose probability). Without normalization, models favor shorter answers unfairly. acc_norm divides by answer length to correct this.
MMLU test results with batch size 16:
Original:
Tasks
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.7763
±
0.0033
- humanities
2
none
acc
↑
0.6948
±
0.0063
- formal_logic
1
none
0
acc
↑
0.5397
±
0.0446
- high_school_european_history
1
none
0
acc
↑
0.8485
±
0.0280
- high_school_us_history
1
none
0
acc
↑
0.9510
±
0.0152
- high_school_world_history
1
none
0
acc
↑
0.9030
±
0.0193
- international_law
1
none
0
acc
↑
0.8926
±
0.0283
- jurisprudence
1
none
0
acc
↑
0.8241
±
0.0368
- logical_fallacies
1
none
0
acc
↑
0.8466
±
0.0283
- moral_disputes
1
none
0
acc
↑
0.8092
±
0.0212
- moral_scenarios
1
none
0
acc
↑
0.4782
±
0.0167
- philosophy
1
none
0
acc
↑
0.8360
±
0.0210
- prehistory
1
none
0
acc
↑
0.8765
±
0.0183
- professional_law
1
none
0
acc
↑
0.5984
±
0.0125
- world_religions
1
none
0
acc
↑
0.8655
±
0.0262
- other
2
none
acc
↑
0.8252
±
0.0065
- business_ethics
1
none
0
acc
↑
0.8100
±
0.0394
- clinical_knowledge
1
none
0
acc
↑
0.8226
±
0.0235
- college_medicine
1
none
0
acc
↑
0.7803
±
0.0316
- global_facts
1
none
0
acc
↑
0.6000
±
0.0492
- human_aging
1
none
0
acc
↑
0.8072
±
0.0265
- management
1
none
0
acc
↑
0.9029
±
0.0293
- marketing
1
none
0
acc
↑
0.9444
±
0.0150
- medical_genetics
1
none
0
acc
↑
0.9000
±
0.0302
- miscellaneous
1
none
0
acc
↑
0.9119
±
0.0101
- nutrition
1
none
0
acc
↑
0.8562
±
0.0201
- professional_accounting
1
none
0
acc
↑
0.6383
±
0.0287
- professional_medicine
1
none
0
acc
↑
0.8603
±
0.0211
- virology
1
none
0
acc
↑
0.5783
±
0.0384
- social sciences
2
none
acc
↑
0.8739
±
0.0059
- econometrics
1
none
0
acc
↑
0.6667
±
0.0443
- high_school_geography
1
none
0
acc
↑
0.9242
±
0.0189
- high_school_government_and_politics
1
none
0
acc
↑
0.9689
±
0.0125
- high_school_macroeconomics
1
none
0
acc
↑
0.8231
±
0.0193
- high_school_microeconomics
1
none
0
acc
↑
0.9160
±
0.0180
- high_school_psychology
1
none
0
acc
↑
0.9413
±
0.0101
- human_sexuality
1
none
0
acc
↑
0.8702
±
0.0295
- professional_psychology
1
none
0
acc
↑
0.8513
±
0.0144
- public_relations
1
none
0
acc
↑
0.8091
±
0.0376
- security_studies
1
none
0
acc
↑
0.8041
±
0.0254
- sociology
1
none
0
acc
↑
0.8905
±
0.0221
- us_foreign_policy
1
none
0
acc
↑
0.9100
±
0.0288
- stem
2
none
acc
↑
0.7545
±
0.0073
- abstract_algebra
1
none
0
acc
↑
0.5600
±
0.0499
- anatomy
1
none
0
acc
↑
0.8519
±
0.0307
- astronomy
1
none
0
acc
↑
0.9079
±
0.0235
- college_biology
1
none
0
acc
↑
0.9306
±
0.0213
- college_chemistry
1
none
0
acc
↑
0.4900
±
0.0502
- college_computer_science
1
none
0
acc
↑
0.6800
±
0.0469
- college_mathematics
1
none
0
acc
↑
0.5200
±
0.0502
- college_physics
1
none
0
acc
↑
0.5784
±
0.0491
- computer_security
1
none
0
acc
↑
0.8400
±
0.0368
- conceptual_physics
1
none
0
acc
↑
0.8426
±
0.0238
- electrical_engineering
1
none
0
acc
↑
0.7793
±
0.0346
- elementary_mathematics
1
none
0
acc
↑
0.7804
±
0.0213
- high_school_biology
1
none
0
acc
↑
0.9226
±
0.0152
- high_school_chemistry
1
none
0
acc
↑
0.7241
±
0.0314
- high_school_computer_science
1
none
0
acc
↑
0.8800
±
0.0327
- high_school_mathematics
1
none
0
acc
↑
0.5815
±
0.0301
- high_school_physics
1
none
0
acc
↑
0.6689
±
0.0384
- high_school_statistics
1
none
0
acc
↑
0.7361
±
0.0301
- machine_learning
1
none
0
acc
↑
0.7143
±
0.0429
Groups
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.7763
±
0.0033
- humanities
2
none
acc
↑
0.6948
±
0.0063
- other
2
none
acc
↑
0.8252
±
0.0065
- social sciences
2
none
acc
↑
0.8739
±
0.0059
- stem
2
none
acc
↑
0.7545
±
0.0073
Heretic:
Tasks
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.7711
±
0.0033
- humanities
2
none
acc
↑
0.6869
±
0.0063
- formal_logic
1
none
0
acc
↑
0.5317
±
0.0446
- high_school_european_history
1
none
0
acc
↑
0.8485
±
0.0280
- high_school_us_history
1
none
0
acc
↑
0.9412
±
0.0165
- high_school_world_history
1
none
0
acc
↑
0.9072
±
0.0189
- international_law
1
none
0
acc
↑
0.8760
±
0.0301
- jurisprudence
1
none
0
acc
↑
0.8426
±
0.0352
- logical_fallacies
1
none
0
acc
↑
0.8221
±
0.0300
- moral_disputes
1
none
0
acc
↑
0.8064
±
0.0213
- moral_scenarios
1
none
0
acc
↑
0.4514
±
0.0166
- philosophy
1
none
0
acc
↑
0.8167
±
0.0220
- prehistory
1
none
0
acc
↑
0.8889
±
0.0175
- professional_law
1
none
0
acc
↑
0.5945
±
0.0125
- world_religions
1
none
0
acc
↑
0.8772
±
0.0252
- other
2
none
acc
↑
0.8230
±
0.0066
- business_ethics
1
none
0
acc
↑
0.8000
±
0.0402
- clinical_knowledge
1
none
0
acc
↑
0.8189
±
0.0237
- college_medicine
1
none
0
acc
↑
0.7688
±
0.0321
- global_facts
1
none
0
acc
↑
0.6300
±
0.0485
- human_aging
1
none
0
acc
↑
0.7937
±
0.0272
- management
1
none
0
acc
↑
0.9126
±
0.0280
- marketing
1
none
0
acc
↑
0.9487
±
0.0145
- medical_genetics
1
none
0
acc
↑
0.8900
±
0.0314
- miscellaneous
1
none
0
acc
↑
0.9055
±
0.0105
- nutrition
1
none
0
acc
↑
0.8497
±
0.0205
- professional_accounting
1
none
0
acc
↑
0.6348
±
0.0287
- professional_medicine
1
none
0
acc
↑
0.8713
±
0.0203
- virology
1
none
0
acc
↑
0.5843
±
0.0384
- social sciences
2
none
acc
↑
0.8684
±
0.0060
- econometrics
1
none
0
acc
↑
0.6579
±
0.0446
- high_school_geography
1
none
0
acc
↑
0.9091
±
0.0205
- high_school_government_and_politics
1
none
0
acc
↑
0.9689
±
0.0125
- high_school_macroeconomics
1
none
0
acc
↑
0.8077
±
0.0200
- high_school_microeconomics
1
none
0
acc
↑
0.9034
±
0.0192
- high_school_psychology
1
none
0
acc
↑
0.9431
±
0.0099
- human_sexuality
1
none
0
acc
↑
0.8550
±
0.0309
- professional_psychology
1
none
0
acc
↑
0.8546
±
0.0143
- public_relations
1
none
0
acc
↑
0.7909
±
0.0390
- security_studies
1
none
0
acc
↑
0.7918
±
0.0260
- sociology
1
none
0
acc
↑
0.8905
±
0.0221
- us_foreign_policy
1
none
0
acc
↑
0.9100
±
0.0288
- stem
2
none
acc
↑
0.7507
±
0.0074
- abstract_algebra
1
none
0
acc
↑
0.5700
±
0.0498
- anatomy
1
none
0
acc
↑
0.8296
±
0.0325
- astronomy
1
none
0
acc
↑
0.8947
±
0.0250
- college_biology
1
none
0
acc
↑
0.9167
±
0.0231
- college_chemistry
1
none
0
acc
↑
0.5200
±
0.0502
- college_computer_science
1
none
0
acc
↑
0.6800
±
0.0469
- college_mathematics
1
none
0
acc
↑
0.5500
±
0.0500
- college_physics
1
none
0
acc
↑
0.6176
±
0.0484
- computer_security
1
none
0
acc
↑
0.8100
±
0.0394
- conceptual_physics
1
none
0
acc
↑
0.8426
±
0.0238
- electrical_engineering
1
none
0
acc
↑
0.7793
±
0.0346
- elementary_mathematics
1
none
0
acc
↑
0.7804
±
0.0213
- high_school_biology
1
none
0
acc
↑
0.9161
±
0.0158
- high_school_chemistry
1
none
0
acc
↑
0.6995
±
0.0323
- high_school_computer_science
1
none
0
acc
↑
0.8800
±
0.0327
- high_school_mathematics
1
none
0
acc
↑
0.5926
±
0.0300
- high_school_physics
1
none
0
acc
↑
0.6623
±
0.0386
- high_school_statistics
1
none
0
acc
↑
0.7083
±
0.0310
- machine_learning
1
none
0
acc
↑
0.6964
±
0.0436
Groups
Version
Filter
n-shot
Metric
Value
Stderr
mmlu
2
none
acc
↑
0.7711
±
0.0033
- humanities
2
none
acc
↑
0.6869
±
0.0063
- other
2
none
acc
↑
0.8230
±
0.0066
- social sciences
2
none
acc
↑
0.8684
±
0.0060
- stem
2
none
acc
↑
0.7507
±
0.0074
MMLU - Massive Multitask Language Understanding, ~14,000 multiple-choice questions across 57 subjects (math, history, law, medicine, etc.).
Upscale redone with the missing final layer included. The original upscales were always missing a layer, but I never troubleshooted to identify *what* layer was missing. Turns out it was the final layer. That's kind of an important one.
This model is an uncensored, creative writing and RP model. Compared to the older version, it is smarter and I think has a bit less repetition. The old V2 version though is slightly more creative due to the instability it had.