News icon

Kimi K3 is now available on Runpod

How to Deploy OpenJev on Runpod Serverless

Deploy OpenJev, a 27B decision model that returns structured answers and confidence scores, on a Runpod Serverless load balancer endpoint, then send your first request with cURL and Python.

How to Deploy OpenJev on Runpod Serverless

Over the past couple of weeks, there has been a lot of chatter about Jev. Jev is a model from TypeSafe AI that outputs structured choices and confidence scores rather than generating natural-language text.

OpenJev is a 27-billion-parameter decision model that runs on a single GPU, and you can deploy it on Runpod Serverless as a load balancer endpoint that bills only while a worker is running. On its authors' 10,000-question text benchmark, the FP8 build this tutorial deploys scored 84.2%, compared with 85.4% for TypeSafe's hosted Jev. OpenJev is an independent project and isn't affiliated with TypeSafe.

A note on licensing: the openjev-worker repository and OpenJev's server code are licensed under Apache 2.0, but the model weights the worker downloads are licensed under CC BY-NC 4.0. That means you can use them for research, learning, and other non-commercial projects, with attribution. For commercial use, the OpenJev team asks you to open a discussion on the Hugging Face repo first.

With a regular LLM, you ask a question and get back text, which you then have to parse. If asked for a star rating, a chatbot might reply "3 stars," "about a 3," or "I'd say 3 out of 5, because...". Your code has to handle all of them. OpenJev, like Jev, works differently. Every question you send includes a type, which tells the model what kind of answer you want, and the answer always comes back in the matching format.

This blog post will walk you through deploying OpenJev on a Runpod Serverless endpoint, sending it your first request, and reading the responses it returns. All of the configuration and code in this post is also in the companion repository, openjev-runpod-serverless, along with a troubleshooting guide. Code from this post can be found in our repository.

Prerequisites

  • A Runpod account. The endpoint runs on an H100 80 GB, and you're billed only while a worker is running. As of September 27, 2026, Runpod's pricing page lists the H100 80 GB flex rate at $4.79/hr, and a cold start that downloads the model uses a few minutes of that before your first request lands. A few dollars covers everything in this tutorial.
  • A Runpod API key. Create one in the Runpod console under Settings and API Keys.
  • Python 3 with the requests library. Install it with pip install requests.
  • A terminal for sending requests and running the example script. cURL is useful for quick tests.
  • A text editor of your choice for editing the Python code sample.
  • Basic familiarity with JSON and HTTP requests. You don't need any machine learning experience.

Deploying an OpenJev endpoint on Serverless

For this example, you will want to use Serverless since it starts a worker only when a request arrives and shuts it down when the endpoint goes idle, so you pay only while a worker is up. That includes the time a worker spends starting and loading the model, not just the time it spends answering requests. See Serverless pricing for the details.

You'll use a load balancer endpoint, since the OpenJev worker is a standard web server that listens for POST requests, just like OpenJev's reference server. The endpoint forwards HTTP requests directly to the server, so the API works as documented, and any client built for the hosted Jev API can use your endpoint by changing its base URL.

OpenJev doesn't come with a ready-made Runpod setup, so you'll want to use openjev-worker, a community-built container image that does the setup for you. It loads the model onto the GPU and runs OpenJev's own API server, so you don't have to install anything yourself.

One important thing to note is that after an idle period, a new worker has to start and load the model, which can take several minutes. A load balancer endpoint doesn't hold your request while that happens: if no worker is ready within about 2 minutes, it returns an error. So the examples below check that a worker is ready before sending a request. You don't need a network volume for this example; the worker downloads the model to the container disk on each cold start. For a production setup, a network volume or Runpod's cached models would keep the weights between cold starts, but that's beyond the scope of this post.

You can deploy in either of two ways: through the console or with your coding agent and Runpod's MCP server.

Deploying using the Console

To get started with OpenJev by using Runpod's console:

  1. Log in to the Runpod Console, select “Serverless” and “New Endpoint,” and press “Next.”
  2. Choose “Deploy from a Docker image” under Custom code and press “Next.”
  3. Configure the image:
    • Container image: ghcr.io/brandonbondig/openjev-worker@sha256:1f95a727a3de6d0a29808f59d89b49424f4d4bc78c871f5080052abb4105424d. This pins the exact version of the image this tutorial was tested with, so a later change to the worker can't break your deployment. The image is public, so you don't need registry authentication or a start command.
    • Expand "Container configuration." Set the container disk to 60 GB (the model weights are about 30 GB) and enter 3000,3001 under "Expose HTTP Ports."
    • Expand "Environment variables" and add PORT = 3000, PORT_HEALTH = 3001, and HEALTH_CHECK_PATH = /.
    • Select "Load balancer" as the endpoint type, not the default "Queue." Click "Next."
    • Why both the variables and the exposed ports: the variables tell the load balancer which ports the API and health check use, and exposing the ports lets traffic reach them. If you skip the exposed ports, the worker loads the model but never receives traffic. Requests return 502 errors, and you pay for about 8 minutes of GPU time before the worker is terminated.
  4. Choose a GPU:
    • Pick the 80 GB H100. The console lists GPUs by memory size rather than by name, with the latest generation shown by default. If you don't see an 80 GB option, click "Show more."
    • Why the H100: the model card's measured setup is a single 80 GB GPU with native FP8 support, and its latency table comes from one H100 running FP8. Other GPUs with FP8 support, such as the L40S or RTX 6000 Ada, have less memory and haven't been measured with OpenJev. An A100 has the memory but not the FP8 support, so vLLM would fall back to a different, unmeasured code path.
  5. Set the worker settings:
    • Max workers: 1. One worker is plenty for this blog post, and it keeps a burst of test requests from starting several GPUs at once.
    • Expand "Advanced settings" for the remaining fields.
    • Active workers: 0, so the endpoint shuts down when idle.
    • Idle timeout: 300 seconds. The default is 5 seconds, which means a worker shuts down almost immediately after each request, and the next request pays for another full model load. Five minutes keeps the worker warm long enough to run the curl example and then the Python script on a single cold start.
    • FlashBoot: leave it on. It's enabled by default under "Performance features," and it helps most when workers cycle between busy and idle. It doesn't speed up the very first boot.
    • You can create a template using these settings. To see an example, take a look at our GitHub.
  6. Click "Deploy." Remember that you're billed from the moment a worker starts, including the minutes it spends loading the model, until it shuts down after the idle timeout.

Deploying using the MCP server

If you prefer a more agentic approach, using our MCP server, you can use a prompt like this one to get started:

Deploy OpenJev on Runpod Serverless using the image ghcr.io/brandonbondig/openjev-worker@sha256:1f95a727a3de6d0a29808f59d89b49424f4d4bc78c871f5080052abb4105424d (use this exact digest, not the latest tag). Create a load-balancer endpoint (not queue-based) named openjev with:

- One H100 80 GB per worker (OpenJev's measured setup is one 80 GB GPU with native FP8, so no A100 fallback). Look up the H100 GPU pool ID from the Runpod catalog rather than guessing it.
- Environment variables PORT=3000, PORT_HEALTH=3001 and HEALTH_CHECK_PATH=/
- A 60 GB container disk
- Expose HTTP ports 3000 and 3001 in the container configuration
- Active workers 0, maximum workers 1, and an idle timeout of 300 seconds
- Request count scaling (load balancer endpoints don't support queue delay scaling)
- FlashBoot on (the API defaults it to off, unlike the console)

Tell me the hourly price before you create anything. When it's done, give me the endpoint URL and a curl command that sends a test request to `/v1/systemone`.

‍

Sending a cURL request

To test the endpoint, you'll ask OpenJev a simple question about a short product review. Imagine you run an online store and want to know what customers think of a pair of headphones. Instead of reading every review yourself, you can ask OpenJev questions about each one and get answers your code can use. For this example, you'll start with a single yes/no question: Does this reviewer like the sound quality?

First, set environment variables for your Runpod API Key and the serverless endpoint ID you previously set up:

‍

export RUNPOD_API_KEY="rpa_yourapikey"
export ENDPOINT_ID="your_endpoint_id"

‍

Next, wait for a worker to be ready. This loop asks the worker for its version. The worker responds to /v1/version once the model is loaded and the API server is up, so a 200 status code indicates it's ready to accept requests. The loop keeps checking for up to 20 minutes. Most failed checks come back quickly with an error body, but if Runpod holds the request while a worker starts, a single check can take about 2 minutes. The loop prints the response body on every failed attempt, because the body tells you whether to keep waiting. "no workers available" is normal: the worker is still starting, so let the loop run. "not allowed for QB API" is not: it means the endpoint was created with the Queue type instead of Load balancer, and no amount of waiting will fix it. Stop the loop and create a new endpoint with the correct type. For other errors, see the troubleshooting table in the companion repository.

ready=0
deadline=$(( $(date +%s) + 1200 ))
while [ "$(date +%s)" -lt "$deadline" ]; do
  resp=$(curl -s -m 130 -w '\n%{http_code}' \
    -H "Authorization: Bearer $RUNPOD_API_KEY" \
    "https://$ENDPOINT_ID.api.runpod.ai/v1/version")
  code=${resp##*$'\n'}
  body=${resp%$'\n'*}
  if [ "$code" = "200" ]; then ready=1; echo "Worker ready."; break; fi
  if [ "$code" = "401" ] || [ "$code" = "403" ]; then echo "HTTP $code: check RUNPOD_API_KEY and its endpoint permissions."; break; fi
  echo "Not ready yet (HTTP $code): ${body:-(no response)}"
  sleep 10
done
[ "$ready" = "1" ] || echo "Worker not ready. Check the messages above and the endpoint's logs before sending requests."

Once the worker is ready, send your request:‍

curl -sS -m 60 https://$ENDPOINT_ID.api.runpod.ai/v1/systemone \
  -H "Authorization: Bearer $RUNPOD_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"openjev","state":"Review: The sound is great, but the battery barely lasts a day.","questions":{"likes_sound":{"type":"noul","instructions":"Does the reviewer like the sound quality?"}}}' \
  -w "\n"

The response you get back should be similar to the following:

{"id": "shim-1790637935267", "model": "openjev-FP8 T=0.85 noul=1.829074,0.0 flags={\"perms\":1,\"stagger\":true,\"loop_break\":false,\"compact\":false,\"compact_cap\":0,\"layout\":\"\",\"pad\":0,\"targeted\":true,\"instr_style\":\"pyrepr\"} shim=shim.py@81a22f1b1b89", "answers": {"likes_sound": {"type": "noul", "noul": 0.9801}}, "usage": {"input_tokens": 75, "output_tokens": 0}}

The reason why this output is in this form is that OpenJev is a specialized model designed to evaluate pre-defined choices or scores and return structured, probabilistic outputs rather than generating conversational text.

The response has four parts:

  • answers holds the answer to each question you asked, under the name you gave it, in this case likes_sound. The type confirms the question type, and noul is the model's confidence that the answer is yes, from 0 to 1. A value of 0.9801 indicates the model is about 98% confident that the reviewer likes the sound.
  • usage shows how much text the model processed. input_tokens counts your review and question (75 tokens here). output_tokens is always 0, because OpenJev doesn't generate text. It returns probabilities directly.
  • id is a unique identifier for this request. It's useful if you're logging requests.
  • model records the model version and the exact settings the server used. You can ignore it for this tutorial. It helps with troubleshooting or checking that two runs used the same configuration.

Using Python to make a request

To make a request using Python, you can use the following script:

import os
import sys
import time

import requests

API_KEY = os.environ.get("RUNPOD_API_KEY")
ENDPOINT_ID = os.environ.get("ENDPOINT_ID")
if not API_KEY or not ENDPOINT_ID:
    sys.exit("Set RUNPOD_API_KEY and ENDPOINT_ID first.")

BASE_URL = f"https://{ENDPOINT_ID}.api.runpod.ai"
HEADERS = {"Authorization": f"Bearer {API_KEY}"}


def wait_until_ready(max_wait=1200):
    """Poll /v1/version until a worker is ready to take requests."""
    deadline = time.time() + max_wait
    while time.time() < deadline:
        try:
            # The worker only answers /v1/version once the model is loaded.
            # If no worker is ready, Runpod returns an error after about
            # 2 minutes, so each check is short and we keep polling.
            r = requests.get(f"{BASE_URL}/v1/version", headers=HEADERS, timeout=130)
            if r.status_code == 200:
                print("Worker ready.")
                return
            if r.status_code in (401, 403):
                sys.exit(f"HTTP {r.status_code}: check RUNPOD_API_KEY and its endpoint permissions.")
            print(f"Not ready yet (HTTP {r.status_code}): {r.text[:200]}")
        except requests.RequestException as e:
            print(f"Not ready yet: {e}")
        time.sleep(10)
    sys.exit("No worker became ready. Check the endpoint's logs in the console.")


def ask(state, questions, attempts=3):
    """Send text plus typed questions to OpenJev and return the answers."""
    body = {"model": "openjev", "state": state, "questions": questions}
    for attempt in range(1, attempts + 1):
        try:
            r = requests.post(
                f"{BASE_URL}/v1/systemone", headers=HEADERS, json=body, timeout=60
            )
            r.raise_for_status()
            return r.json()["answers"]
        except requests.RequestException as e:
            # Covers timeouts, HTTP errors, and dropped connections. If the
            # model server fails, it closes the connection with no JSON body,
            # which raises ConnectionError rather than Timeout.
            print(f"Attempt {attempt} failed: {e}")
            if attempt < attempts:
                time.sleep(10)
    sys.exit(f"No answer after {attempts} attempts. Check the endpoint's logs.")


review = (
    "Review: I bought these headphones for my commute. The sound is great "
    "and the noise cancelling works well on the train, but the battery "
    "barely lasts a day and they hurt my ears after an hour."
)

questions = {
    "sentiment": {
        "type": "choice",
        "instructions": "What is the overall tone of this review?",
        "criteria": {"positive": None, "mixed": None, "negative": None},
    },
    "recommends": {
        "type": "noul",
        "instructions": "Does the reviewer say they would recommend this product?",
    },
    "rating": {
        "type": "score",
        "instructions": "How many stars would this reviewer likely give?",
        "criteria": ["1 star", "2 stars", "3 stars", "4 stars", "5 stars"],
    },
}

wait_until_ready()
answers = ask(review, questions)

# choice: the option the model picked
sentiment = answers["sentiment"]
print(f"Sentiment: {sentiment['choice']} (confidence {sentiment['confidence']})")

# noul: the probability that the answer is yes
p_yes = answers["recommends"]["noul"]
print(f"Recommends: {'yes' if p_yes >= 0.5 else 'no'} (probability of yes {p_yes})")

# score: a position on the scale, counted from 0; legend maps positions to labels
rating = answers["rating"]
label = rating["legend"][str(round(rating["score"]))]
print(f"Rating: {label} (score {rating['score']}, confidence {rating['confidence']})")

Save the script in the editor of your choosing, and in your terminal, you can run the script:

python openjev_runpod.py

The response you get back should look similar to the following:

Worker ready.
Sentiment: mixed (confidence 0.9871)
Recommends: no (probability of yes 0.061)
Rating: 3 stars (score 1.9857, confidence 0.9274)

The Python script sends the same kind of request as the curl command: a POST to /v1/systemone with a review and some questions. The difference is what happens around it. Instead of one yes/no question, the script asks three questions of different types in a single request, then reads each answer and prints it in plain English, such as turning a score of 1.9857 into "3 stars." It also handles the cold start for you, waiting until a worker is ready before it sends the request, and retrying if the connection drops, which makes it closer to what you'd use in a real application. curl is good for a quick check that the endpoint works, and Python is how you'd build something on top of it.

Next steps

You now have OpenJev running on a Runpod Serverless endpoint that starts a worker when a request arrives and shuts it down when it goes idle. You've sent it a review and a set of typed questions, and you've seen that every answer comes back as a probability in the format you asked for, with no text to parse. The worker's request and response shapes follow the hosted Jev API, so a client written for hosted Jev can point at your endpoint by changing its base URL. The model runs entirely on your endpoint; nothing is forwarded to TypeSafe. Be sure to check out the code, template, and examples in our repository.

This is just the start of what you can do with OpenJev and Runpod Serverless. As a next step, you could extend the script to analyze a batch of reviews or flag content for a closer look, whether for a research project, a personal tool, or a prototype. If you're thinking about a commercial product, check the license first. The weights are CC BY-NC 4.0, so you'll need permission from the OpenJev authors. Be sure to let us know in the #built-on-runpod channel on Discord if this post inspires you to build anything.

Build what’s next.

Build, train, and scale AI workloads on Runpod with cloud GPUs, Serverless, and Clusters.

Star field background