Skip to main content
← All field notes

Comparison / On-Device AI

OBLITERATUS MMLU gaps: what conditions hide

Joshua Hrisko, Principal Engineer at MakerPortal

Joshua HriskoPrincipal Engineer

6 min readSan Francisco, CA

OBLITERATUS MMLU gaps: what conditions hide
AI-generated illustration · decorative; every figure in this post comes from the sources cited below

Composed from the signals scanned on 2026-08-21.

The −6.0pp MMLU capability loss attributed to V1 and the −0.3pp attributed to V2 were produced under sample sizes that differ by a factor of ten: n=285 for V1, n=2,850 for V2 and the stock baseline. The V1 figure has a further problem. 84.60 − 81.4 = 3.2pp, not 6.0pp — the arithmetic from the numbers OBLITERATUS supplies does not reproduce the claimed gap, which means the stock baseline used for V1’s evaluation was not the 84.60% figure appearing in the same README table.

Figures side by side

MMLU accuracy and measurement conditions

ModelMMLU scoreToolShotn (questions)Stated gap vs stock
Stock Qwen3.8-27B84.60%lm-eval0-shot2,850
OBLITERATED V284.3%lm-eval0-shot2,850−0.3pp
OBLITERATED V181.4%lm-eval0-shot285−6.0pp (stated; see below)

The V2 delta is computed against 84.60% at the same n on both sides. For the V1 delta to reach −6.0pp, the stock reference at n=285 would have to be 87.4% (81.4 + 6.0 = 87.4) — a figure absent from the supplied data. OBLITERATUS reports no stock evaluation at n=285.

Refusal rate on the same 842-prompt corpus

ModelCorpus sizeRefusalsRateHow measured
Stock Qwen3.8-27Bnot evaluated on corpus~100% (estimate)Not counted
OBLITERATED V184200.00%Pass/fail count
OBLITERATED V284220.24%Pass/fail count

The README states the stock figure as an approximation, not a count against the 842-item corpus; only the V1 and V2 rows share a common measurement procedure.

MMLU subcategory shift, V2 vs stock (both at n=2,850)

SubcategoryV2 vs stock
Humanities+1.4pp
Social sciences−1.5pp
All other subcategories (derived)−0.2pp net

The overall delta is −0.3pp. Humanities (+1.4pp) and social sciences (−1.5pp) net to −0.1pp, so the remaining −0.2pp is distributed across subcategories the README does not individually break out.

Ship score (internal metric only)

ModelShip score
OBLITERATED V188.7
OBLITERATED V292.1

The README does not define ship score’s composition or computation method. Both variants were evaluated by the same scorer, making the +3.4 improvement internally consistent — but without a scorer definition, the figure cannot be anchored to any external benchmark from the supplied information.

Where the comparison holds and where it stops

The V2-vs-stock MMLU comparison is the one clean number here. Both runs used lm-eval at 0-shot with n=2,850; −0.3pp is a like-for-like measurement. OBLITERATUS reports a 95% confidence interval of ±0.65pp at that sample size, putting the observed delta within noise at that significance level. The subcategory breakdown is also measured under matching conditions and worth tracking independently: abliteration surgeries act on direction vectors in the residual stream, and uneven redistribution across knowledge domains — humanities recovering +1.4pp while social sciences loses −1.5pp — is mechanistically plausible rather than incidental.

The V1/V2 refusal comparison is also clean: same corpus, same pass/fail criterion, both raw counts. V1 is strictly more permissive (0/842 vs 2/842). The README presents those two refusals as the deliberate price of MMLU recovery through blending rather than single-surgery abliteration — a framing the data supports.

That claim — that V2 represents a “massive improvement over V1’s −6.0pp gap” — runs into two compounding problems.

First, the sample sizes differ by a factor of ten. At n=285 (5 questions per subject across 57 subjects), a single-question error on one subject moves its contribution by 20 percentage points — noise that propagates into the aggregate. The −6.0pp and −0.3pp figures were not produced under equivalent conditions; their arithmetic difference does not directly measure the improvement from V1 to V2.

Second, the arithmetic is inconsistent. The README supplies stock = 84.60% and V1 = 81.4%. Their difference is 3.2pp. Recovering the −6.0pp figure requires a stock score of 87.4% at n=285 — roughly 2.8pp above the n=2,850 stock figure. Plausible under sampling variance at small n, but the figure is not reported and cannot be verified from what OBLITERATUS supplies.

The temperature and system-prompt guidance in the README — that the model states greedy decoding “produces the most complete, code-rich outputs” and that “temps above 0.5 degrade quality significantly” — are stated as conclusions from A/B testing with no sample sizes or scoring criteria attached. Observed author preference, not a reproducible measurement.

Reported by OBLITERATUS/Qwen3.8-27B-OBLITER… GB
safetensors total size 54 GB Q8_0 GGUF size 27 GB Q6_K GGUF size 21 GB Q5_K_M GGUF size 18 GB Q4_K_M GGUF size 16 GB

Chart: figures as OBLITERATUS/Qwen3.8-27B-OBLITERATED reports them, drawn from the quantities this post cites. Bars are proportional to the reported values; the studio has not re-measured them.

What cannot be concluded

The 842-prompt refusal corpus is not described in enough detail to assess adversarial coverage. Two refusals in 842 attempts could represent a narrow edge-case failure or a systematic residual — the count alone does not distinguish. The stock “~100%” estimate is not a count against those same 842 prompts, so the absolute reduction from stock to either abliterated variant cannot be stated as a precise figure; both approach zero on this corpus, and that is all the data supports. Whether either generalises to prompt distributions outside these 842 items is a separate question the supplied data does not address. For memory-allocation considerations when running the GGUF variants (ranging from ~14 GB for IQ4_XS to ~27 GB for Q8_0) on constrained hardware, our serial loading note addresses reclamation strategies that are orthogonal to the capability figures here.


Method: this note was drafted by us.anthropic.claude-sonnet-4-6 from a single source — the model card for OBLITERATUS/Qwen3.8-27B-OBLITERATED. Before publication an automated gate re-checked every extracted claim against the source document (13 claim(s) and 19 quantity(ies) verified) and every claim-shaped number in the draft against that evidence (5 derived from it, no rows from our own tables were supplied to the draft). The studio has not re-run OBLITERATUS/Qwen3.8-27B-OBLITERATED’s benchmarks; figures attributed to it are its own.

FAQ

Is the −0.3pp MMLU delta actually within noise?

OBLITERATUS reports a 95% confidence interval of ±0.65pp at n=2,850, and the observed −0.3pp falls inside it — consistent with zero capability loss at that significance level. Whether the true underlying difference is zero requires more than a single evaluation run to establish. The supplied measurement does not contradict the zero-loss claim.

Why doesn't 84.60 − 81.4 equal the claimed −6.0pp for V1?

It equals 3.2pp. For the gap to be 6.0pp, the stock model would have had to score 87.4% on the n=285 subsample used for V1's evaluation. The README reports no stock run at n=285, so the reference point for the stated −6.0pp is missing from the supplied data. The V1-to-V2 improvement in blending technique is plausible on other grounds, but the specific magnitude — recovering 5.7pp out of 6.0pp — cannot be verified from what is published.

Does V2's 0.24% refusal rate make it less useful than V1 for bypass use cases?

On the 842-prompt corpus, V1 refused zero and V2 refused two. V2 is marginally less aggressive by this single measure. The README attributes this directly to the 60/40 blend retaining more safety-adjacent circuitry than V1's single aggressive surgery. Whether a difference of two prompts in 842 constitutes an operational problem depends entirely on the target prompt distribution, not these counts.

What is "ship score" and can it be compared to anything external?

The README does not define ship score's composition or the tooling behind it. OBLITERATUS reports V1 at 88.7 and V2 at 92.1; the direction is consistent with the stated goal of capability preservation. Without a scorer definition, neither figure can be reproduced or mapped to any external evaluation. Until OBLITERATUS publishes the scorer, treat it as an opaque internal summary with no external anchor.

Recommended Studio & Hardware Gear

Affiliate links support independent R&D

Tested studio equipment and reference hardware utilized for this build. Product images & pricing sourced from Amazon Creators API / SparkFun Electronics.