One Ruler for Everyone: The Cyseclens Rubric for Re-Scoring MITRE Evaluations
Before I score a single vendor, here’s exactly how the scoring works. Poke holes in it.
One Ruler for Everyone: The Cyseclens Rubric for Re-Scoring MITRE Evaluations
Before I score a single vendor, here’s exactly how the scoring works. Poke holes in it.
Last week I argued that the MITRE ATT&CK Evaluations have a strange property: they’re rigorous and objective, yet nearly every vendor walks away claiming victory. That happens because MITRE publishes the evidence but deliberately assigns no score — so each vendor invents its own.
I said I’d fix that by applying one transparent rubric to every vendor. And I promised to publish the rubric before any results, so the method can be judged on its own merits and nobody can accuse me of reverse-engineering a scoring system to make a predetermined villain look bad.
This is that rubric. It’s version 1. I want you to try to break it.
The one rule that governs everything
Every choice below serves a single principle: the same rules apply to every vendor, defined in advance, with no exceptions for anyone. A vendor doesn’t get a better mapping because it’s popular, or a worse one because it’s an easy target. The rubric is fixed; the results fall where they fall.
What MITRE actually records
For each step of a simulated attack, MITRE records how — or whether — each product surfaced the activity. The detection categories have evolved across rounds, but they map onto a consistent ladder of usefulness to a defender, from most to least actionable:
- Technique — the product not only flagged the activity but identified the specific ATT&CK technique behind it. This is the most actionable outcome: the analyst is told what is happening and how.
- Tactic — the product identified the adversary’s goal (the “why”) but not the specific technique.
- General — the product flagged the activity as suspicious or malicious, but with no ATT&CK context. Something’s wrong, but you’re on your own to figure out what.
- Telemetry — the data was captured and is available, but the product didn’t raise anything. An analyst who goes looking can find it; the product won’t tell them to look.
- None — the activity wasn’t captured or surfaced at all.
MITRE also flags two important modifiers: whether a detection required a configuration change made during the evaluation, and whether the detection was delayed (for example, held for sandbox detonation or human analysis before surfacing).
The rubric
Step 1 — Base credit, by detection quality. Each in-scope substep earns credit according to the most useful detection the product produced for it:
- Technique — 1.0
- Tactic — 0.7
- General — 0.4
- Telemetry — 0.2
- None — 0.0
Step 2 — Apply modifiers, as multipliers. These reduce the base credit to reflect real-world usefulness:
- Required a configuration change during the test — × 0.5
- Detection was delayed — × 0.75
Modifiers stack. A technique-level detection that required a config change and arrived delayed earns 1.0 × 0.5 × 0.75 = 0.375. Final credit is capped to the 0–1 range.
Step 3 — Aggregate to a vendor score. A vendor’s score for a round is the average credit across all in-scope substeps, expressed as a percentage. Substeps MITRE marks Not Applicable are excluded from the denominator entirely — they’re not counted for or against anyone. Substeps where the result was None are counted, as zeros; a miss is part of the record.
Step 4 — The claims gap. For each vendor I record, with an archived source, the headline number that vendor marketed about the same round. The Cyseclens number minus the vendor’s marketed number is the claims gap — the distance between the story and the evidence.
The judgment calls, defended
A rubric is only honest if it owns its choices. Here are the three that matter most, and why I made them. These are exactly where I expect — and want — disagreement.
Telemetry earns credit, but little. Telemetry is not nothing: a product that captures the data lets a skilled analyst hunt, even if it never alerts. But it is also not a detection — it’s logging that puts the entire burden on the human. Scoring it at 0.2 says “this has real value, but treating it as equivalent to an alert that names the technique is how marketing math inflates scores.” If you think telemetry deserves more, or nothing at all, that’s a fair fight — tell me why.
Configuration changes are penalized, not excluded. When a product only catches an attack step after its configuration was changed mid-evaluation, that tells you something real: out of the box, it missed. But excluding those detections entirely would over-punish — the capability does exist, and a well-run team might have configured it that way from day one. Halving the credit tries to hold both truths: it counts, but not as much as something that worked without intervention.
Delayed detections are discounted, not dropped. A detection that arrives after the attack step has completed is still evidence the product saw it — but operationally, a late alert may arrive after the damage is done. A 0.75 multiplier reflects “better late than never, but timing matters.”
What this rubric deliberately does not measure
Honesty about scope is part of the method. This rubric scores detection depth and quality in the MITRE evaluation, and nothing else. It does not measure:
- False positives / alert fatigue — MITRE’s evaluation isn’t designed to surface these, and pretending otherwise would be dishonest.
- Protection (whether the product blocked the attack, as opposed to detecting it) — a separate MITRE result that deserves its own treatment, not a blend that hides which is which.
- Price, usability, support, or scale — all real buying factors, none of them measurable from this dataset.
A single number is a starting point for a conversation, not a verdict on a product. Anyone who tells you one score settles a six-figure decision is selling something.
A worked example
Suppose a product faced five substeps and produced: Technique; Telemetry; Technique but only after a configuration change; None; Tactic (delayed).
- Technique → 1.0
- Telemetry → 0.2
- Technique × config change (0.5) → 0.5
- None → 0.0
- Tactic (0.7) × delayed (0.75) → 0.525
Average = (1.0 + 0.2 + 0.5 + 0.0 + 0.525) / 5 = 44.5%.
If that same vendor’s marketing announced, say, “100% detection coverage” for the round — because something was captured on four of five steps — the claims gap is the distance between 100% and 44.5%. That distance is the entire point of this project.
Now break it
This is version 1, and it’s public on purpose. I’d rather someone finds a flaw now, before I score anyone, than after. So:
- Do the weights look wrong? Say so, and say what they should be.
- Is there a modifier I’m missing, or one I shouldn’t penalize?
- Is my treatment of telemetry too generous, or too harsh?
Tell me in the comments or by message. Substantive critiques will change the rubric — and I’ll changelog every version publicly, so the method has a paper trail of its own. When I map this to a specific evaluation round, I’ll publish the exact category-by-category mapping for that round alongside the results, so every number is traceable to MITRE’s own data.
The rubric is set. Next comes the part everyone’s actually waiting for: applying it, to everyone, and publishing the gaps.
Which round or category should I run first? I’m listening.
Cyseclens is an independent project analyzing cybersecurity vendors using only public evidence. One transparent standard, applied to every vendor. No sponsorships. Every number traces to a source you can check. Follow for the first re-score.
메타데이터
- post_id
- e73c0bc6b067
- slug
- one-ruler-for-everyone-the-cyseclens-rubric-for-re-scoring-mitre-evaluations-e73c0bc6b067
- url
- https://medium.com/cyseclens/one-ruler-for-everyone-the-cyseclens-rubric-for-re-scoring-mitre-evaluations-e73c0bc6b067
- canonical_url
- https://medium.com/cyseclens/one-ruler-for-everyone-the-cyseclens-rubric-for-re-scoring-mitre-evaluations-e73c0bc6b067
- author_url
- https://medium.com/@miteshpant
- status
- ok
- fetched_at
- 2026-08-15 18:16:42