Running jev-lite: A Calibrated Decision Model on Your Own 4090

How to rent a GPU on lium.io and self-host vagmi/jev-lite, an open-source Gemma adapter that answers typed questions with calibrated probabilities instead of free text.

24 September 20263 min read

Most "LLM classifiers" are a prompt, a generation, and a json.loads() you're praying doesn't throw. jev-lite skips all of that. It's a small QLoRA adapter on top of Gemma 4 E4B that reads a situation, reads a typed question about it, and hands back a calibrated probability distribution in a single forward pass — no generation, no parsing, no "as an AI language model." It's 35M trainable parameters on an 8B base, it runs on one RTX 4090, and it's fully open. Here's how to stand it up yourself and what you get back.

What it actually answers

jev-lite speaks three question types:

  • noul — is this true? Returns one probability.
  • choice — pick from named options. Returns a probability per option plus a confidence.
  • score — rate against ordered levels. Returns an expected level, not just a bucket.

Because the answer is read straight off the logits at the option letters, the model can't answer outside the options you gave it. There's no retry loop because there's nothing to retry.

Renting the GPU

I used lium.io to rent a spot 4090 by the hour instead of buying one. The CLI:

curl -fsSL https://lium.io/install.sh | bash
lium init                     # paste your API key
lium ls --gpu RTX4090 --format json     # find something cheap and available
lium up <index> --name jevlite --yes    # rent it

Once the pod is running, check the Ports tab — Lium maps an internal container port to an external one it exposes on your pod's public IP, so whatever you bind inside the container isn't necessarily the port you hit from outside:

INTERNAL   EXTERNAL
22         50009
8888       50008

Bind your server to the internal port; hit the pod at <public_ip>:<external_port>.

Setup and serve, in two commands

The model card points at a small reference server — vagmi/jevlite — that speaks a clean wire API (POST /v1/systemone) over either a 4-bit torch backend or vLLM. On the pod:

python3 run.py setup
#!/usr/bin/env python3
REPO = "https://github.com/vagmi/jevlite.git"
SETUP_CMD = (
    "curl -LsSf https://astral.sh/uv/install.sh | sh && "
    "export PATH=$HOME/.local/bin:$PATH && "
    f"(test -d jevlite || git clone {REPO}) && "
    "cd jevlite && uv sync"
)
SERVE_CMD = (
    "cd jevlite && export PATH=$HOME/.local/bin:$PATH && "
    "uv run python serve.py --host 0.0.0.0 --port 8888"
)

Then, after exporting a Hugging Face token that's accepted the Gemma license (the base model is gated):

export HF_TOKEN=hf_xxx
python3 run.py serve

First run downloads google/gemma-4-E4B-it (4-bit NF4) plus the jev-lite adapter from the Hub, loads them onto the GPU, and starts serving. No separate download step — the server does it on first boot.

Hitting it from your laptop

curl -s http://<pod_public_ip>:50008/v1/systemone \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": "This is the third time this month my order has arrived damaged. I am done giving you my money.",
    "questions": {
      "frustration": {
        "type": "score",
        "instructions": "How frustrated is the customer?",
        "criteria": ["Calm", "Annoyed", "Frustrated", "Very angry"]
      },
      "will_churn": {
        "type": "noul",
        "instructions": "Is this customer at risk of churning?"
      }
    }
  }'
{
  "answers": {
    "frustration": {
      "score": 2.6629,
      "legend": {"0": "Calm", "1": "Annoyed", "2": "Frustrated", "3": "Very angry"},
      "probabilities": {"0": 0.0114, "1": 0.0214, "2": 0.2601, "3": 0.7071},
      "confidence": 0.7071
    },
    "will_churn": { "noul": 0.9707 }
  }
}

The score type doesn't just pick a bucket — it puts a probability on every level and reports the expected value (2.66, between "Frustrated" and "Very angry"), with confidence telling you how much of that mass actually sits on the rounded level. will_churn at 0.97 is about as unambiguous as this stuff gets.

The part that actually matters: it knows when it doesn't know

Here's the same server on a genuinely split case — a failing-payouts ticket that could plausibly go to either billing or technical:

{
  "choice": "billing",
  "probabilities": { "billing": 0.5622, "technical": 0.4378 },
  "confidence": 0.5622
}

And on an unambiguous one — a happy user asking a how-to question:

{
  "choice": "product",
  "probabilities": { "billing": 0.0073, "technical": 0.0328, "product": 0.9598 },
  "confidence": 0.9598
}

Same model, same wire format — but 0.56 vs 0.96 confidence tells you exactly which of these you can auto-route and which needs a human glance. That gap is the entire point of the adapter: on held-out data it hits 81.6% accuracy with an Expected Calibration Error of 0.019 — when it says 80% confident, it's right about 80% of the time. That's what lets you build a gate: route the confident 62% of traffic automatically, send the 0.5–0.8 band to review, and kick anything under 0.5 straight to a human.

Why this instead of a bigger model + a prompt

You could ask GPT-whatever to "classify this and return JSON," but then you're parsing free text, hoping it didn't wrap the answer in a sentence, and getting back a confidence number the model made up because you asked for one — not one that's been measured against reality. jev-lite skips generation entirely, reads the answer off the logits, and its confidence number is calibrated against 1,898 held-out rows. It also happens to be small enough and open enough that "self-host it on a spot GPU for a few dollars an hour" is a completely reasonable answer to "how do I run this."

One caveat worth flagging before you point this at production traffic: the adapter's training mix includes ANLI and RACE, both non-commercial-licensed datasets, so treat jev-lite itself as research/non-commercial use unless you retrain it without those sources — the repo's build_data.py takes a --sets flag to do exactly that.