Running jev-lite: A Calibrated Decision Model on Your Own 4090
How to rent a GPU on lium.io and self-host vagmi/jev-lite, an open-source Gemma adapter that answers typed questions with calibrated probabilities instead of free text.
Most "LLM classifiers" are a prompt, a generation, and a json.loads() you're
praying doesn't throw. jev-lite skips
all of that. It's a small QLoRA adapter on top of Gemma 4 E4B that reads a
situation, reads a typed question about it, and hands back a calibrated
probability distribution in a single forward pass — no generation, no parsing,
no "as an AI language model." It's 35M trainable parameters on an 8B base, it
runs on one RTX 4090, and it's fully open. Here's how to stand it up yourself
and what you get back.
What it actually answers
jev-lite speaks three question types:
noul— is this true? Returns one probability.choice— pick from named options. Returns a probability per option plus a confidence.score— rate against ordered levels. Returns an expected level, not just a bucket.
Because the answer is read straight off the logits at the option letters, the model can't answer outside the options you gave it. There's no retry loop because there's nothing to retry.
Renting the GPU
I used lium.io to rent a spot 4090 by the hour instead of buying one. The CLI:
curl -fsSL https://lium.io/install.sh | bash
lium init # paste your API key
lium ls --gpu RTX4090 --format json # find something cheap and available
lium up <index> --name jevlite --yes # rent it
Once the pod is running, check the Ports tab — Lium maps an internal container port to an external one it exposes on your pod's public IP, so whatever you bind inside the container isn't necessarily the port you hit from outside:
INTERNAL EXTERNAL
22 50009
8888 50008
Bind your server to the internal port; hit the pod at <public_ip>:<external_port>.
Setup and serve, in two commands
The model card points at a small reference server —
vagmi/jevlite — that speaks a clean wire
API (POST /v1/systemone) over either a 4-bit torch backend or vLLM. On the
pod:
python3 run.py setup
#!/usr/bin/env python3
REPO = "https://github.com/vagmi/jevlite.git"
SETUP_CMD = (
"curl -LsSf https://astral.sh/uv/install.sh | sh && "
"export PATH=$HOME/.local/bin:$PATH && "
f"(test -d jevlite || git clone {REPO}) && "
"cd jevlite && uv sync"
)
SERVE_CMD = (
"cd jevlite && export PATH=$HOME/.local/bin:$PATH && "
"uv run python serve.py --host 0.0.0.0 --port 8888"
)
Then, after exporting a Hugging Face token that's accepted the Gemma license (the base model is gated):
export HF_TOKEN=hf_xxx
python3 run.py serve
First run downloads google/gemma-4-E4B-it (4-bit NF4) plus the jev-lite
adapter from the Hub, loads them onto the GPU, and starts serving. No separate
download step — the server does it on first boot.
Hitting it from your laptop
curl -s http://<pod_public_ip>:50008/v1/systemone \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "This is the third time this month my order has arrived damaged. I am done giving you my money.",
"questions": {
"frustration": {
"type": "score",
"instructions": "How frustrated is the customer?",
"criteria": ["Calm", "Annoyed", "Frustrated", "Very angry"]
},
"will_churn": {
"type": "noul",
"instructions": "Is this customer at risk of churning?"
}
}
}'
{
"answers": {
"frustration": {
"score": 2.6629,
"legend": {"0": "Calm", "1": "Annoyed", "2": "Frustrated", "3": "Very angry"},
"probabilities": {"0": 0.0114, "1": 0.0214, "2": 0.2601, "3": 0.7071},
"confidence": 0.7071
},
"will_churn": { "noul": 0.9707 }
}
}
The score type doesn't just pick a bucket — it puts a probability on every
level and reports the expected value (2.66, between "Frustrated" and "Very
angry"), with confidence telling you how much of that mass actually sits on
the rounded level. will_churn at 0.97 is about as unambiguous as this stuff
gets.
The part that actually matters: it knows when it doesn't know
Here's the same server on a genuinely split case — a failing-payouts ticket that could plausibly go to either billing or technical:
{
"choice": "billing",
"probabilities": { "billing": 0.5622, "technical": 0.4378 },
"confidence": 0.5622
}
And on an unambiguous one — a happy user asking a how-to question:
{
"choice": "product",
"probabilities": { "billing": 0.0073, "technical": 0.0328, "product": 0.9598 },
"confidence": 0.9598
}
Same model, same wire format — but 0.56 vs 0.96 confidence tells you
exactly which of these you can auto-route and which needs a human glance. That
gap is the entire point of the adapter: on held-out data it hits 81.6%
accuracy with an Expected Calibration Error of 0.019 — when it says 80%
confident, it's right about 80% of the time. That's what lets you build a
gate: route the confident 62% of traffic automatically, send the 0.5–0.8 band
to review, and kick anything under 0.5 straight to a human.
Why this instead of a bigger model + a prompt
You could ask GPT-whatever to "classify this and return JSON," but then you're
parsing free text, hoping it didn't wrap the answer in a sentence, and getting
back a confidence number the model made up because you asked for one — not
one that's been measured against reality. jev-lite skips generation
entirely, reads the answer off the logits, and its confidence number is
calibrated against 1,898 held-out rows. It also happens to be small enough and
open enough that "self-host it on a spot GPU for a few dollars an hour" is a
completely reasonable answer to "how do I run this."
One caveat worth flagging before you point this at production traffic: the
adapter's training mix includes ANLI and RACE, both non-commercial-licensed
datasets, so treat jev-lite itself as research/non-commercial use unless
you retrain it without those sources — the repo's build_data.py takes a
--sets flag to do exactly that.