← Back to list

Who Verifies the Verifier?

Building a Verifier-Guided RAG + SLM Learning Loop with QLoRA and Jev

AISHWARYA P.S.V.S · 2026-10-02 19:31 · 0 claps · 15.2 min read
#jevs #qlora-fine-tuning #rags #lms #fine-tuning
Open on Medium ↗
Wiki topics: RAG · RAG & Retrieval FT · Fine-tuning & Adaptation EDU · Education & Learning

Who Verifies the Verifier?

Building a Verifier-Guided RAG + SLM Learning Loop with QLoRA and Jev

Most Retrieval-Augmented Generation systems end with a familiar architecture:

User Question
     │
     ▼
Retriever
     │
     ▼
Relevant Context
     │
     ▼
LLM / SLM
     │
     ▼
Answer

But I wanted to explore what happens after that.

What if the retriever finds the correct evidence, but the model still interprets it incorrectly?

What if a Small Language Model drops an important condition such as:

only if

or:

unless

What if we add a verifier?

And then an even more interesting question appears:

What if the verifier is also wrong?

That question became the most important part of this experiment.

I built an end-to-end system combining:

  • ContractNLI
  • Qwen2.5–0.5B-Instruct
  • teacher-generated supervision
  • human ground-truth labels
  • Jev
  • QLoRA
  • Retrieval-Augmented Generation
  • claim-level verification
  • deterministic routing
  • held-out benchmarking
  • hard-example mining

The objective was not to claim state-of-the-art performance.

The goal was to understand whether verification could become part of the training, inference, evaluation, and future learning loop of a small language model.

The Core Problem

Small Language Models are increasingly attractive because they can be:

  • cheaper to run
  • easier to deploy locally
  • faster for constrained workloads
  • easier to fine-tune
  • practical on limited hardware

But smaller models can also be less reliable.

A particularly dangerous class of errors happens when retrieval works correctly, but generation does not preserve the meaning of the retrieved evidence.

Consider this evidence:

Opened headphones can be returned only if they are defective.

Now imagine the model responds:

Yes, you can return opened headphones.

At first glance, this sounds reasonable.

But the model dropped the critical phrase:

only if they are defective

That tiny omission completely changes the policy.

The problem is therefore not always:

Retriever failed

Sometimes the problem is:

Retriever succeeded
        │
        ▼
Correct evidence reached the model
        │
        ▼
Model interpreted it incorrectly

This kind of failure matters in areas such as:

  • contracts
  • finance
  • enterprise policy
  • compliance
  • legal workflows
  • customer support
  • healthcare systems

So I wanted another layer between generation and the final answer.

That layer became Jev.

Why I Chose Jev

One of my design principles was to separate:

Generation
    from
Judgment

The student SLM is responsible for generating language.

Jev is responsible for making a focused decision.

Conceptually:

┌───────────────────┐
        │ Retrieved Evidence│
        └─────────┬─────────┘
                  │
                  ▼
        ┌───────────────────┐
        │ Generated Claim   │
        └─────────┬─────────┘
                  │
                  ▼
        ┌───────────────────┐
        │        Jev        │
        └─────────┬─────────┘
                  │
       ┌──────────┼───────────┐
       │          │           │
       ▼          ▼           ▼
 SUPPORTED   CONTRADICTED  INSUFFICIENT
       │          │           │
       ▼          ▼           ▼
    RETURN     ESCALATE   RETRIEVE MORE

This produces a useful separation of responsibilities:

SLM
 │
 └──► Generate
Jev
 │
 └──► Judge
Router
 │
 └──► Act

The verifier is therefore not responsible for writing another long answer.

Its output becomes a machine-actionable decision.

Where Jev Fits in the Architecture

I used Jev in two different stages.

Training-time verification

Teacher Prediction
        │
        ▼
Compare With Human Label
        │
        ▼
Jev Independently Checks
Original Context + Claim
        │
        ▼
Validated Training Example

Inference-time verification

Student Answer
      │
      ▼
Split Into Claims
      │
      ▼
Jev Checks Each Claim
Against Retrieved Evidence
      │
      ▼
Routing Decision

This distinction is important.

Jev is:

NOT the teacher
NOT the human label
NOT automatically ground truth

It is an independent decision component.

That became extremely important later when I benchmarked the system.

Full Architecture

The system evolved into three connected loops:

  1. Training
  2. Production
  3. Improvement

End-to-End Architecture

TRAINING LOOP
              ┌────────────────────┐
              │    ContractNLI     │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Context + Claim    │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │    Teacher LLM     │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Verdict + Rationale│
              │ + Grounded Answer  │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Compare With Human │
              │ Ground-Truth Label │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │  Jev Verification  │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Validated Examples │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │       QLoRA        │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │   Student SLM v1   │
              └──────────┬─────────┘
                   PRODUCTION LOOP
                         │
                         ▼
              ┌────────────────────┐
              │   User Question    │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │   RAG Retrieval    │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Relevant Evidence  │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Student SLM Draft  │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │  Split Into Claims │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Jev Claim Checks   │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Aggregate Routing  │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Grounded Response  │
              └──────────┬─────────┘
                  IMPROVEMENT LOOP
                         │
                         ▼
              ┌────────────────────┐
              │   Failure Logs     │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Failure Validation │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │ Trusted Hard Cases │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │      QLoRA v2      │
              └──────────┬─────────┘
                         │
                         ▼
              ┌────────────────────┐
              │   Student SLM v2   │
              └────────────────────┘

Step 1: Starting With ContractNLI

I used ContractNLI because its examples naturally contain the kind of reasoning I wanted to test.

Each example contains:

Contract Context
      +
Claim
      +
Human Label

The classification labels are:

SUPPORTED
CONTRADICTED
INSUFFICIENT

For example:

Contract:
The agreement may be terminated with 30 days written notice.
Claim:
The agreement can be terminated without notice.

Ground truth:

CONTRADICTED

The dataset contains many examples involving:

  • conditions
  • obligations
  • exceptions
  • missing evidence
  • contradictions

Exactly the sort of cases where a small model may incorrectly infer more than the evidence actually says.

Step 2: The First Important Failure — Label Leakage

One of the most useful lessons came from a mistake in my first implementation.

Initially, the human label accidentally appeared inside the teacher prompt.

That produced responses such as:

The human label states SUPPORTED...

The outputs looked reasonable.

But scientifically, this was wrong.

The teacher already knew the answer.

This is label leakage.

So I changed the architecture.

The teacher now sees only:

Context
   +
Claim

The human label remains hidden.

The teacher is required to independently generate:

Verdict:
Rationale:
Grounded response:

The flow changed from:

Human Label
     │
     ▼
Teacher Sees Answer
     │
     ▼
Writes Explanation

to:

Context + Claim
      │
      ▼
Teacher Predicts Independently
      │
      ▼
Prediction Generated
      │
      ▼
Human Label Revealed
Only For Evaluation

This was a crucial correction.

Step 3: Teacher-Generated Supervision

The teacher produces structured outputs such as:

Verdict: SUPPORTED
Rationale:
The contract explicitly states that the recipient receives
no rights to the confidential information.
Grounded response:
The agreement grants the recipient no rights to the
confidential information.

This response becomes potential supervision for the student.

This is not logit-based distillation.

The experiment uses something closer to:

Teacher Response Generation
          │
          ▼
Validated Sequence Targets
          │
          ▼
Supervised QLoRA Fine-Tuning

So the student learns from:

  • verdict format
  • rationale structure
  • grounded responses

Step 4: Why Teacher Outputs Needed Verification

Synthetic supervision can contain errors.

If the teacher generates the wrong answer and I train the student on it, the error simply moves from one model to another.

So I added multiple signals.

Context + Claim
                     /          \
                    /            \
                   ▼              ▼
          ┌──────────────┐ ┌──────────────┐
          │   Teacher    │ │     Jev      │
          └──────┬───────┘ └──────┬───────┘
                 │                │
                 ▼                ▼
          Teacher Verdict     Jev Verdict
                 \                /
                  \              /
                   ▼            ▼
                  Human Ground Truth

I now had three pieces of information:

1. Human Label
2. Teacher Verdict
3. Jev Verdict

The goal was to prevent obviously unreliable examples from entering the student training set.

Step 5: What Happened on the First 20 Examples?

I ran an initial smoke test on 20 examples.

The result:

Generated: 20
Accepted:  10
Rejected:  10

A frequent failure looked like:

Teacher: SUPPORTED
Human:   INSUFFICIENT

The tiny teacher tended to infer support even when the contract did not explicitly establish the claim.

That was already an interesting observation.

It reinforced one important principle:

Synthetic data should not automatically be trusted.

Step 6: QLoRA Distillation

For the student, I used:

Qwen/Qwen2.5-0.5B-Instruct

The setup was:

Base Model
   │
   ▼
4-bit Quantization
   │
   +
   │
LoRA Adapters
   │
   +
   │
Validated Training Data
   │
   ▼
QLoRA Training

The smoke-test training used:

10 validated examples
3 epochs
9 optimization steps

The resulting adapter was saved as:

artifacts/student-qlora-v1

The reported final training loss was approximately:

2.19

This result did not prove that the student had become generally better.

Ten examples are far too few for that.

At this point, the goal was simply:

Can the full pipeline train successfully?

The answer was yes.

Step 7: Connecting the Student to RAG

The student was then placed inside a RAG pipeline.

User Question
      │
      ▼
┌──────────────────┐
│    Retriever     │
└────────┬─────────┘
         │
         ▼
┌──────────────────┐
│ Relevant Evidence│
└────────┬─────────┘
         │
         ▼
┌──────────────────┐
│ Distilled SLM v1 │
└────────┬─────────┘
         │
         ▼
┌──────────────────┐
│   Draft Answer   │
└──────────────────┘

This is where the experiment became much more interesting.

The Headphones Example

I tested the system with:

Question:
Can I return opened headphones?

The retriever returned:

Opened headphones can be returned only if they are defective.

So retrieval had succeeded.

But the SLM produced an answer effectively saying:

Yes, you can return opened headphones.

It also suggested that there was no specific condition.

But the evidence clearly said:

only if they are defective

The model had removed the condition.

This created a very useful diagnosis:

Retriever
    │
    ▼
Correct Evidence
    │
    │ ✅
    ▼
Student SLM
    │
    ▼
Condition Dropped
    │
    │ ❌
    ▼
Incorrect Interpretation

The problem was not retrieval.

It was generation.

Step 8: Jev Caught the Hallucination

Rather than returning the entire draft directly to the user, I split it into individual claims.

Student Draft
                      │
          ┌───────────┼───────────┐
          │           │           │
          ▼           ▼           ▼
       Claim 1     Claim 2     Claim 3
          │           │           │
          ▼           ▼           ▼
         Jev         Jev         Jev

For the claim:

Yes, you can return opened headphones.

Jev returned:

CONTRADICTED

Other generated claims were classified as:

CONTRADICTED

or:

INSUFFICIENT

So in this concrete example:

Jev successfully caught the SLM’s unsupported interpretation.

The complete failure-detection flow looked like this:

┌─────────────────────────────────────────┐
│ Evidence                                │
│                                         │
│ "Opened headphones can be returned     │
│  only if they are defective."          │
└───────────────────┬─────────────────────┘
                    │
                    ▼
          ┌──────────────────┐
          │ Student SLM      │
          └────────┬─────────┘
                   │
                   ▼
┌─────────────────────────────────────────┐
│ Generated Claim                         │
│                                         │
│ "Yes, you can return opened headphones."│
└───────────────────┬─────────────────────┘
                    │
                    ▼
             ┌─────────────┐
             │     Jev     │
             └──────┬──────┘
                    │
                    ▼
             CONTRADICTED
                    │
                    ▼
                ESCALATE
                    │
                    ▼
┌─────────────────────────────────────────┐
│ Grounded Response                       │
│                                         │
│ Opened headphones can be returned       │
│ only if they are defective.             │
└─────────────────────────────────────────┘

This was one of the strongest moments in the project.

The SLM made a semantic error.

Jev caught it.

And the bad draft was not returned as the final answer.

Step 9: Claim-Level Routing

The verifier itself is not the final product.

Its decisions need to affect application behavior.

The basic router became:

Jev Verdict
                        │
        ┌───────────────┼────────────────┐
        │               │                │
        ▼               ▼                ▼
   SUPPORTED       CONTRADICTED      INSUFFICIENT
        │               │                │
        ▼               ▼                ▼
     RETURN          ESCALATE        RETRIEVE MORE

This is one of the reasons I found the verifier useful.

Its result could directly control normal application code.

Step 10: Why Aggregation Was Necessary

Initially, every claim produced an independent action.

For example:

Claim 1 → ESCALATE
Claim 2 → ESCALATE
Claim 3 → RETRIEVE
Claim 4 → ESCALATE
Claim 5 → RETRIEVE

That was useful for debugging.

But it was not a good user experience.

So I added an aggregate routing rule:

ESCALATE
   >
RETRIEVE
   >
RETURN

Conceptually:

Claim-Level Decisions
       ESCALATE
       ESCALATE
       RETRIEVE
       ESCALATE
       RETRIEVE
           │
           ▼
   ┌──────────────────┐
   │ Aggregate Router │
   └────────┬─────────┘
            │
            ▼
         ESCALATE
            │
            ▼
   One Final Grounded
        Response

For the headphones example:

RETURN:    0
RETRIEVE:  1
ESCALATE:  5

The final answer became:

Based on the retrieved evidence:
Opened headphones can be returned only if they are defective.

At This Point, the System Looked Successful

I had a nice story:

Correct Evidence
      │
      ▼
SLM Drops Condition
      │
      ▼
Jev Detects Contradiction
      │
      ▼
Router Blocks Draft
      │
      ▼
Grounded Answer Returned

It would have been easy to stop here and say:

“Adding a verifier solved the hallucination problem.”

But one successful example does not prove that.

So I benchmarked the system.

That changed the story.

Step 11: Held-Out Evaluation

I created another 20 ContractNLI examples that did not overlap with the original smoke-test rows.

Then I compared:

┌──────────────────────────────┐
│        Held-Out Data         │
└──────────────┬───────────────┘
               │
     ┌─────────┼──────────┐
     │         │          │
     ▼         ▼          ▼
 Base SLM   SLM v1   SLM v1 + Jev

A future experiment will also include:

Hard-example-trained SLM v2

Benchmark Results

The results were:

SystemAccuracyMacro-F1Base SLM20%0.148Distilled SLM v120%0.142

This result immediately showed:

The first QLoRA run did not improve held-out classification accuracy.

That is not surprising.

The student had only 10 validated training examples.

The fine-tuning run was a smoke test.

But the important point is that I now had measured evidence.

Step 12: Evaluating the Jev Gate

The Jev-gated version produced:

Total Examples: 20
Returned:        6
Escalated:      14
Coverage:       30%

Of the six predictions the system allowed through:

Correct Returned:   2
Incorrect Returned: 4

Therefore:

Accepted Accuracy:
33.3%
False Acceptance Rate:
66.7%

Visualizing that:

20 Held-Out Cases
                        │
                        ▼
              ┌─────────────────┐
              │ SLM + Jev Gate  │
              └────────┬────────┘
                       │
          ┌────────────┴────────────┐
          │                         │
          ▼                         ▼
     6 Returned                14 Escalated
          │
      ┌───┴────┐
      │        │
      ▼        ▼
   2 Correct  4 Wrong

That changed my interpretation of the architecture.

Who Verifies the Verifier?

Originally, I assumed:

Student Prediction
      +
Jev Prediction
      │
      ▼
Both Agree
      │
      ▼
Probably Safe

The benchmark showed that this assumption was too strong.

Sometimes:

Student
   │
   └──► Wrong
Jev
   │
   └──► Same Wrong Decision

The two systems could agree and still be wrong.

That became the central lesson of the project.

A simplistic architecture is:

Generator
    │
    ▼
Verifier
    │
    ▼
Truth

A more realistic architecture is:

Generator
    │
    ▼
Verifier
    │
    ▼
Verifier Evaluation
    │
    ▼
Calibration
    │
    ▼
Routing Policy
    │
    ▼
Final Action

A verifier is another model component.

It also needs:

  • benchmarking
  • calibration
  • monitoring
  • error analysis
  • stress testing

Did Jev Help?

Yes.

But the answer is more nuanced than:

Jev = always correct

Jev helped in four important ways.

1. It caught a real hallucination

The headphones example clearly showed a case where the SLM dropped a critical condition and Jev marked the resulting claim as contradicted.

2. It exposed uncertainty

On the held-out benchmark:

70% of examples were escalated

instead of being blindly returned.

That itself is useful information.

3. It provided structured actions

SUPPORTED
→ RETURN
CONTRADICTED
→ ESCALATE
INSUFFICIENT
→ RETRIEVE MORE

4. It made the pipeline observable

Without a verifier:

Model
  │
  ▼
Answer
  │
  ▼
User

With a verifier:

Model
  │
  ▼
Claim
  │
  ▼
Verifier
  │
  ▼
Decision
  │
  ▼
Router
  │
  ▼
Logs + Metrics
  │
  ▼
Final Response

That gives the system measurable internal behavior.

Why I Did Not Immediately Train SLM v2

My original idea was:

SLM v1
  │
  ▼
Jev Finds Failure
  │
  ▼
Save Failure
  │
  ▼
QLoRA
  │
  ▼
SLM v2

But the benchmark revealed a risk.

If the verifier itself can be wrong:

Incorrect Jev Decision
        │
        ▼
Incorrect Hard Example
        │
        ▼
QLoRA
        │
        ▼
Student Learns
Incorrect Behavior

That would turn verification errors into training errors.

So the improvement loop needs another validation step.

A Safer Hard-Example Loop

The updated design is:

Student Failure
                       │
                       ▼
                 Jev Signal
                       │
                       ▼
             ┌─────────────────┐
             │ Validate Failure│
             └────────┬────────┘
                      │
         ┌────────────┼────────────┐
         │            │            │
         ▼            ▼            ▼
    Human Label   Trusted LLM   Rule Check
         │            │            │
         └────────────┼────────────┘
                      │
                      ▼
              Corrected Target
                      │
                      ▼
              Trusted Hard Example
                      │
                      ▼
                  QLoRA v2

This is a much safer self-improvement mechanism.

Error Taxonomy

Not every failure should become a training example.

So I want to classify failures into categories such as:

Dropped Condition
Missed Negation
Incorrect Exception
Unsupported Assumption
Wrong Entity
Wrong Date
Wrong Amount
Retrieval Failure
Verifier Disagreement

Why?

Because different problems require different fixes.

For example:

Retrieval Failed
      │
      ▼
Better Chunking
Better Embeddings
Query Rewriting
Reranking

But:

Correct Evidence Retrieved
          +
Student Dropped "only if"
          │
          ▼
Good Fine-Tuning Candidate

This distinction matters.

Condition Preservation

The headphones example highlighted a particularly important family of expressions:

only if
unless
except
provided that
subject to
within 30 days
not permitted unless

These words often carry more meaning than the rest of the sentence.

For example:

Refund allowed only if defective.

is very different from:

Refund allowed.

So one metric I want to add is:

Condition Preservation Accuracy

Example test cases:

Refund allowed only if defective.
Cancellation allowed unless processing has started.
Access permitted except for external contractors.
Disclosure permitted provided that written approval exists.

This kind of metric may be particularly useful for enterprise and policy applications.

Adaptive Retrieval

Verification can also improve retrieval.

Instead of:

Retrieve Top-3
     │
     ▼
Generate
     │
     ▼
INSUFFICIENT
     │
     ▼
Stop

the system can do:

Retrieve Top-3
     │
     ▼
Generate
     │
     ▼
Jev: INSUFFICIENT
     │
     ▼
Retrieve More
     │
     ▼
Top-8
     │
     ▼
Rerank
     │
     ▼
Generate Again

Now verification becomes part of an active control loop.

Query Decomposition

Complex questions can also be decomposed.

Complex Question
                    │
                    ▼
            ┌──────────────┐
            │  Decomposer  │
            └──────┬───────┘
                   │
      ┌────────────┼────────────┐
      │            │            │
      ▼            ▼            ▼
Sub-question 1 Sub-question 2 Sub-question 3
      │            │            │
      ▼            ▼            ▼
   Retrieve      Retrieve      Retrieve
      │            │            │
      └────────────┼────────────┘
                   │
                   ▼
             Merge Evidence
                   │
                   ▼
                Generate

This can improve evidence coverage before the student even begins generation.

Structured Student Output

Another improvement would be to make the student produce structured results.

Instead of:

Opened headphones can be returned...

the student could return:

{
  "verdict": "INSUFFICIENT",
  "conditions": [
    "Return is allowed only if defective"
  ],
  "answer": "Opened headphones can be returned only if they are defective."
}

Then I can evaluate independently:

Verdict Accuracy
Condition Preservation
Final Answer Groundedness

Confidence Calibration

A verifier confidence score should not automatically be interpreted as:

Probability that this answer is correct

Instead, confidence needs calibration.

For example:

0.9–1.0 confidence
      │
      ▼
How often actually correct?
0.8–0.9 confidence
      │
      ▼
How often actually correct?
0.7–0.8 confidence
      │
      ▼
How often actually correct?

Possible methods include:

Reliability Diagrams
Expected Calibration Error
Platt Scaling
Isotonic Regression

Only then should confidence thresholds control important production decisions.

Multi-Verifier Architecture

Another possible extension is to avoid trusting one verifier.

Claim
                           │
             ┌─────────────┼─────────────┐
             │             │             │
             ▼             ▼             ▼
            Jev        NLI Model     Rule Engine
             │             │             │
             └─────────────┼─────────────┘
                           │
                           ▼
                  Decision Aggregator
                           │
                  ┌────────┴─────────┐
                  │                  │
                  ▼                  ▼
               RETURN            ESCALATE

For example:

Jev:
SUPPORTED
NLI:
CONTRADICTED
Result:
ESCALATE

The idea is not to stack models endlessly.

The goal is to see whether independent verification signals can meaningfully reduce false acceptance.

Future SLM v2

Once I have trustworthy hard examples, the next training stage becomes:

Clean Validated Data
                   │
                   │
                   +
                   │
           Trusted Hard Cases
                   │
                   ▼
             Mixed Dataset
                   │
                   ▼
                 QLoRA
                   │
                   ▼
              Student v2

Then I can run the full comparison:

Held-Out Dataset
       │
       ├──► Base SLM
       │
       ├──► Distilled SLM v1
       │
       ├──► Distilled SLM v1 + Jev
       │
       └──► Hard-Example SLM v2

Metrics should include:

Accuracy
Macro-F1
Precision
Recall
Coverage
Escalation Rate
False Acceptance Rate
Condition Preservation
Latency
Verifier Agreement

Only then can I make a stronger claim about model improvement.

Champion vs Challenger

A production system should not automatically deploy the newest adapter.

Instead:

┌─────────────────┐
│ Student SLM v1  │
│    Champion     │
└────────┬────────┘
         │
         │ compare
         │
┌────────▼────────┐
│ Student SLM v2  │
│   Challenger    │
└─────────────────┘

Promote v2 only if:

Accuracy improves
        AND
False acceptance decreases
        AND
Latency stays acceptable

Otherwise:

Keep v1

Training completion is not the same thing as production readiness.

How Is This Different From Typical Jev + RAG Work?

A common Jev + RAG architecture looks like:

Query
  │
  ▼
Retriever
  │
  ▼
Many Chunks
  │
  ▼
Jev Reranking / Filtering
  │
  ▼
Best Chunks
  │
  ▼
LLM
  │
  ▼
Answer

That is a useful use case.

Jev can also be used for:

Classification
Routing
Relevance Scoring
Citation Checking
Document Filtering
Structured Decisions

My experiment explores a broader lifecycle.

TRAINING
                        │
                        ▼
Human Ground Truth
        │
        ▼
Teacher
        │
        ▼
Jev Validation
        │
        ▼
QLoRA Student
        │
        │
        ▼
                    PRODUCTION
                        │
                        ▼
RAG
 │
 ▼
Student Generation
 │
 ▼
Claim-Level Jev Verification
 │
 ▼
Routing
 │
 ▼
Grounded Response
 │
 │
 ▼
                    EVALUATION
                        │
                        ▼
Held-Out Benchmark
 │
 ▼
Verifier Failure Analysis
 │
 ▼
Trusted Hard Examples
 │
 ▼
Future SLM v2

The differentiator is not simply:

"I used Jev."

It is the integration of:

Distillation
      +
Verification
      +
RAG
      +
Routing
      +
Benchmarking
      +
Failure Analysis
      +
Future Retraining

into one experiment.

What Surprised Me Most

I initially expected the main result to be:

Jev caught an SLM hallucination.

And it did.

But the more interesting result was:

Jev itself could not simply be treated as ground truth.

That changed the design.

My first mental model was:

SLM Makes Mistake
      │
      ▼
Verifier Catches Mistake
      │
      ▼
Problem Solved

The architecture I ended up with is:

SLM Makes Mistake
      │
      ▼
Verifier Judges It
      │
      ▼
Measure Verifier
      │
      ▼
Calibrate
      │
      ▼
Route Carefully
      │
      ▼
Validate Failures
      │
      ▼
Create Trusted Training Data
      │
      ▼
Retrain

That is a much more realistic engineering loop.

What This Experiment Does Not Prove

With such a small initial experiment, I do not claim that:

The distilled model is generally better.
Jev is always correct.
Ten examples are enough for meaningful fine-tuning.
The system is already self-improving.
The current architecture is optimal.

The held-out benchmark actually showed that:

Base Accuracy:
20%
Distilled v1 Accuracy:
20%

So v1 did not outperform the base model.

That is part of the result.

What the Experiment Does Demonstrate

The project successfully connects:

Human Ground Truth
        │
        ▼
Teacher Generation
        │
        ▼
Training-Time Verification
        │
        ▼
QLoRA
        │
        ▼
Student SLM
        │
        ▼
RAG
        │
        ▼
Claim-Level Verification
        │
        ▼
Routing
        │
        ▼
Held-Out Benchmarking
        │
        ▼
Failure Analysis

into one functioning experimental pipeline.

More importantly, the benchmark exposed weaknesses that a demo-only project would have hidden.

Final Takeaway

I started the project with this question:

Can a verifier make a Small Language Model more reliable?

The experiment led me to a more interesting question:

How do we know when the verifier itself is reliable enough to trust?

The solution is probably not one magical model.

It is a system built around:

Generation
    +
Verification
    +
Measurement
    +
Calibration
    +
Routing
    +
Feedback

The direction is not simply:

Train a better model.

It is:

Build a system
      │
      ▼
Detect where the model is weak
      │
      ▼
Measure the weakness
      │
      ▼
Route unsafe cases
      │
      ▼
Validate failures
      │
      ▼
Turn trustworthy failures
into better training data

That is the direction I want to continue exploring.

Not a perfect model.

A system that learns where it is weak.

Final System Vision

┌──────────────┐
                     │    DATA      │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │   TEACHER    │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │ GROUND TRUTH │
                     │  VALIDATION  │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │     JEV      │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │ CLEAN DATA   │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │    QLoRA     │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │ STUDENT SLM  │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │     RAG      │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │    DRAFT     │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │ CLAIM SPLIT  │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │     JEV      │
                     │ VERIFICATION │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │   ROUTING    │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │ FINAL ANSWER │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │ FAILURE LOG  │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │ VALIDATE HARD│
                     │   EXAMPLES   │
                     └──────┬───────┘
                            │
                            ▼
                     ┌──────────────┐
                     │   QLoRA v2   │
                     └──────────────┘

A verifier should not merely catch failures. A well-designed system should measure the verifier, learn from trustworthy failures, and gradually convert those failures into better training data.

References

  1. ContractNLI: A Dataset for Document-Level Natural Language Inference for Contracts Koreeda, Y., Manning, C. D. https://aclanthology.org/2021.findings-emnlp.164/
  2. QLoRA: Efficient Finetuning of Quantized LLMs Dettmers, T. et al. https://arxiv.org/abs/2305.14314
  3. Qwen2.5 Technical Report Qwen Team. https://arxiv.org/abs/2412.15115
  4. TypeSafe AI — Jev Official Jev overview and structured decision workflow. https://www.typesafeai.org/jev
  5. Spring AI + TypeSafe Jev for Modular RAG Example of Jev used for RAG filtering and reranking. https://spring.io/blog/2026/10/02/spring-ai-modular-rag-typesafe-jev/
  6. PEFT: Parameter-Efficient Fine-Tuning Hugging Face documentation. https://huggingface.co/docs/peft/

Project Repository

The complete implementation for this experiment is available on GitHub:

JevRagSlmImprovement https://github.com/aishwaryatesting/JevRagSlmImprovement

My Linkeldn Profile: https://www.linkedin.com/in/aishwarya-p-s-v-s-5a07a41ab/?isSelfProfile=true


메타데이터
post_id
bb77e8eea1b7
slug
who-verifies-the-verifier-bb77e8eea1b7
url
https://medium.com/@psvsaishwarya/who-verifies-the-verifier-bb77e8eea1b7
canonical_url
https://medium.com/@psvsaishwarya/who-verifies-the-verifier-bb77e8eea1b7
author_url
https://medium.com/@psvsaishwarya
status
ok
fetched_at
2026-10-03 09:03:59