25 min read

Install and Self-Host jevos for Local AI Decisions

Install jevos, verify its local yes/no inference API, and securely connect the working HTTP service to remote applications with Localtonet.

A remote application reaches a locally hosted jevos service through a Localtonet tunnel.
jevos performs yes-or-no inference locally while Localtonet provides the remote connection path.
AI and Machine Learning ยท jevos ยท Localtonet ยท 2026

Run focused yes-or-no AI inference locally, validate the API, then connect it to remote applications

jevos is a self-hosted inference service for answering yes-or-no questions about text or structured JSON input. In this guide, we install the project using its documented uv workflow, download its GGUF model and machine-specific llama.cpp runtime, start the local HTTP server, and verify its health and inference endpoints. We also explain request design, routine operation, command-line inference, and common failure modes. Once the service works locally, we show how to expose it through an HTTP tunnel with Localtonet without configuring inbound router port forwarding, a public IP address, or a VPN.

๐Ÿ”’ Local-first inference and controlled remote exposure ๐ŸŒ Documented HTTP API on 127.0.0.1:8017 โšก CPU-oriented yes-or-no classification

What jevos does and when to use it

jevos is an open-source, local inference project designed for binary decisions. An application sends some state, asks one or more yes-or-no questions, and receives a probability for each answer. Instead of asking a general-purpose model to produce prose, the caller gets a value named noul between 0 and 1. That value represents the model's probability that the answer is yes.

This structure is useful for classification, triage, routing, moderation assistance, and policy evaluation. For example, an application can ask whether a support message describes a billing problem, whether a customer appears upset, or whether a request satisfies a policy stated in the question. The application remains responsible for deciding what probability threshold to use and what action to take.

The documented server implements the TypeSafe Jev wire format for yes-or-no questions. Its primary endpoint is POST /v1/systemone. The project also documents GET /v1/models, GET /health, and a browser demo at /dino. By default, the service listens on 127.0.0.1:8017, which limits direct access to the same machine.

๐Ÿง  Binary inference Each supported question returns a probability from 0 to 1 indicating how strongly the model favors yes.
๐Ÿ“ฆ Structured state The request state can be a string, JSON object, or JSON array, allowing callers to classify both natural language and application records.
โ“ Multiple questions One request can contain several named questions that share the same state, so the input state is read once.
๐Ÿ’ป Local CPU execution The documented quickstart explicitly starts the GGUF model on a CPU and lets the operator choose the thread count.
๐Ÿ”Œ HTTP integration Applications can call the local service through JSON HTTP requests or use a Python HTTP client such as requests.
๐Ÿ“„ File-based decisions The jev decide command can process a request file without keeping the HTTP server running.

jevos is intentionally narrower than a general chat service. The documented interface supports yes-or-no questions. Choice and score question types are refused with HTTP status 422. The optional criteria field from the compatible wire format is accepted but is not read, so decision rules should be placed directly in the question's instructions.

A probability is not a business decision by itself

A value above or below 0.5 can be a convenient example threshold, but the correct threshold depends on the consequences of false positives and false negatives. Evaluate the model on representative data before allowing its output to trigger consequential actions.

Prerequisites and documented limitations

The project quickstart assumes that you have a local copy of the repository, the uv command is available, and you can download the model release asset. It then uses uv sync to prepare the project environment. The supplied project documentation does not establish an operating-system support matrix, a specific Python version requirement, a supported uv installation command, or minimum RAM and disk requirements.

Before continuing, make sure you have the following:

  • A local checkout or downloaded copy of the jev repository.
  • A working installation of uv.
  • Enough local storage for the repository, dependencies, runtime, and jevos-q4_k_m.gguf model.
  • Permission to download project dependencies and the machine-specific llama.cpp runtime.
  • A CPU thread count appropriate for your machine and other workloads.
  • curl or another HTTP client for local API verification.
  • The Localtonet client later, if remote access is required.

Run the project commands from the repository directory containing its pyproject.toml and lock file. The documented quickstart does not provide a Docker installation, container image, operating-system service definition, reverse-proxy configuration, or production orchestration architecture. This guide therefore follows the native uv workflow rather than inventing an unsupported container deployment.

Requirement What is documented What you must determine locally
Project environment The quickstart uses uv sync. How to install uv on your operating system.
Model Use the jevos-q4_k_m.gguf release asset. Where to store it if you do not place it in the repository directory.
Inference runtime jev download --only runtime downloads llama.cpp for the current machine. Whether the downloaded runtime is compatible with your particular host.
CPU allocation The server defaults to 4 threads; the example uses 16. A thread count that leaves sufficient capacity for other applications.
Network binding The default host and port are 127.0.0.1 and 8017. Whether another process already uses that port.
Authentication and production deployment These are not established by the supplied README evidence. The access-control and deployment layer appropriate for your environment.
Do not assume the API is safe for unrestricted public access

The supplied project documentation does not establish built-in authentication, authorization, TLS configuration, request limits, or a production hardening model. Keep the initial server bound to 127.0.0.1, test locally, and define an authentication and authorization strategy before sharing the service with untrusted users.

Install and start jevos

The following sequence preserves the project's documented quickstart. Commands assume your terminal is open in the repository directory. If your model is stored elsewhere, replace the filename with its real local path. Do not copy a placeholder path that does not exist.

1

Download the documented GGUF model

Download the jevos-q4_k_m.gguf asset from the project's jevos release. Place it in the repository directory if you want to use the exact relative filename shown by the quickstart. Confirm that the completed file is named jevos-q4_k_m.gguf rather than a partial download or an automatically renamed duplicate.

2

Synchronize the project environment

From the repository directory, run uv sync. This prepares the environment from the project's declared dependency and lock information.

3

Download the machine-specific runtime

Run the documented download command with --only runtime. The project describes this as downloading llama.cpp for the current machine.

4

Start the local HTTP server

Start jevos with the downloaded GGUF model, select the CPU explicitly, and choose an appropriate thread count. The documented example uses 16 threads. The server option default is 4, and the project recommends using fewer threads if other heavy applications are running.

Run the synchronization command:

uv sync

Download the runtime:

uv run jev download --only runtime

Start the server using the documented model and CPU configuration:

uv run jev serve --gguf jevos-q4_k_m.gguf --device cpu --threads 16

Keep this terminal running. Unless you override the server options, the service listens on 127.0.0.1:8017. A process bound to 127.0.0.1 accepts connections originating from the same host, not direct connections addressed to the machine from another device.

Choosing a CPU thread count

The documented default is 4 threads, while the quickstart demonstrates 16. More threads are not automatically the best choice for every system. If the machine also runs a web application, database, desktop workloads, or other inference processes, allocate fewer threads so that jevos does not contend for all available CPU capacity.

Start with a value supported by your actual hardware, submit representative requests, and observe responsiveness across the whole machine. The project documentation does not provide a universal sizing formula, so capacity planning should be based on your workload rather than copied from the example.

Changing the host or port

The server documents --host and --port options, with defaults of 127.0.0.1 and 8017. For this Localtonet workflow, changing the host is normally unnecessary because our client can run on the same device and connect to the loopback service.

Keeping the loopback binding also avoids opening the application directly to the local network. If you change the port, use the same value consistently in local tests, application configuration, and the Localtonet tunnel target.

Verify the local server before remote access

A local HTTP request reaches the running jevos server and returns a yes result.
Test the jevos HTTP service locally before creating remote access.

Do not create a tunnel until the application responds locally. Local verification separates application problems from tunnel configuration problems and gives you a known-good target.

Check readiness

The health endpoint returns a response containing {"status":"ready", ...} once the model has loaded. Call it from the same machine:

curl http://127.0.0.1:8017/health

A connection error means the HTTP process is not reachable at that address. Check whether the server terminal is still running, whether model loading failed, and whether you changed the host or port. A response that is not ready may indicate that startup is still in progress.

Inspect available model names

The models endpoint returns the served model and its jev-latest alias:

curl http://127.0.0.1:8017/v1/models

Using jev-latest lets a client address the served model through the compatible alias. The API also accepts the served model's name, and the documentation states that any jev-* name works for this field.

Submit a yes-or-no decision

Send a JSON body to POST /v1/systemone. This example asks whether a customer statement describes a billing problem:

curl http://127.0.0.1:8017/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "jev-latest",
    "state": "I was charged twice for the same order.",
    "questions": {
      "billing": {
        "type": "noul",
        "instructions": "Is this a billing problem?"
      }
    }
  }'

A successful response contains the model name, an answers object keyed by your question name, and usage information. The relevant value is found at answers.billing.noul. It will be between 0 and 1.

{
  "model": "jevos-q4_k_m",
  "answers": {
    "billing": {
      "type": "noul",
      "noul": 0.9
    }
  },
  "usage": {
    "input_tokens": 27,
    "output_tokens": 0
  }
}

Model output is probabilistic, so do not expect every installation or request variation to reproduce the example probability exactly. Verify the response shape and range, then test behavior with examples from your intended domain.

Open the demonstration page

With the server running locally, open http://127.0.0.1:8017/dino in a browser on the same machine. This is useful as a quick interactive check, but API clients should integrate with the documented JSON endpoints rather than depend on the demo interface.

Inspect inference timing separately from total request time

Each response includes a Server-Timing header carrying inference time. End-to-end latency observed by a client can also include connection setup, HTTP processing, tunnel transit, application code, and queueing under load.

Design reliable yes-or-no requests

A precise binary question and local context flow through jevos to a yes or no result.
Reliable requests provide relevant context and ask one explicit binary question.

The quality of a binary decision depends heavily on the state and question. The API does not turn an ambiguous policy into an unambiguous rule. Give the model the relevant facts, ask one clearly defined yes-or-no question, and include decision rules in instructions.

Use the state for evidence

The state field may be a string, object, or array. A string works for a message or document excerpt. An object is often better when the application already has structured facts because it preserves field boundaries.

{
  "model": "jev-latest",
  "state": {
    "item": "wireless mouse",
    "delivered": "5 days ago",
    "customer_message": "The box arrived empty. This is the second time!"
  },
  "questions": {
    "refund": {
      "type": "noul",
      "instructions": "Our policy refunds items reported missing within 30 days of delivery. Should this customer get a refund?"
    },
    "upset": {
      "type": "noul",
      "instructions": "Is the customer upset?"
    },
    "wrong_item": {
      "type": "noul",
      "instructions": "Does the customer say they received the wrong item?"
    }
  }
}

Questions in the same request share the state. The project reports that shared-state requests can be more efficient than making a separate request for every question because the state is read once. Group questions when they genuinely evaluate the same evidence, but keep names stable so downstream code can reliably map each answer.

Put the rule in the instructions

If a decision depends on a refund window, eligibility rule, severity definition, or another policy, state that rule in the relevant question. Do not rely on the model to know a private organizational policy. The optional criteria field is accepted but not read, so placing the decisive rule there would not produce the intended behavior.

Interpret probabilities deliberately

A simple integration might classify values over 0.5 as yes. That is only a starting point. A workflow where a false positive merely adds a label can tolerate a different threshold from one where a false positive automatically issues a refund or blocks a user.

Build an evaluation set containing clear positives, clear negatives, ambiguous cases, incomplete records, adversarial wording, and examples near policy boundaries. Record the returned probability, compare it with the expected decision, and choose thresholds based on acceptable error tradeoffs. For sensitive outcomes, route uncertain probabilities to human review instead of forcing every input into an automated action.

Call the API from Python

The following documented pattern uses the Python requests package and reads the named probability:

import requests

answer = requests.post(
    "http://127.0.0.1:8017/v1/systemone",
    json={
        "model": "jev-latest",
        "state": "I was charged twice for the same order.",
        "questions": {
            "billing": {
                "type": "noul",
                "instructions": "Is this a billing problem?"
            }
        },
    },
).json()

if answer["answers"]["billing"]["noul"] > 0.5:
    print("send to billing")

Production application code should also handle connection failures, timeouts, non-success HTTP statuses, malformed responses, and missing keys. Those safeguards are application responsibilities and are not shown in the minimal project example.

Local operation and file-based inference

Keep the server command running for applications that need an HTTP endpoint. Stopping that process stops the local API, and a Localtonet tunnel cannot make an inactive target available. If the machine reboots, the supplied evidence does not establish an automatic startup mechanism, so restart behavior must be designed for your operating environment.

The documentation does not provide a supported systemd unit, Windows service definition, macOS launch configuration, process supervisor setup, or container health policy. Avoid copying an unverified service file into a production host. A human reviewer or administrator should define startup, restart, logging, user permissions, and resource limits according to the target system.

Run a single decision without a server

For batch jobs or local scripts, jev decide processes a request file using the same body shape as POST /v1/systemone:

uv run jev decide --gguf jevos-q4_k_m.gguf --device cpu request.json

By default, the answer is printed. To write it to a new file, use the documented output option:

uv run jev decide --gguf jevos-q4_k_m.gguf --device cpu request.json --output answer.json

An existing output file is never overwritten. Choose a new filename, archive or rename the old output, or remove it deliberately after confirming it is no longer needed.

If the request file omits model, the command treats it as the engine's native request and returns the engine's fuller output, including probabilities, prompt hashes, and timings. Use the HTTP-compatible body when you want the same request and response shape across file-based and server-based workflows.

Mode Best suited to Operational behavior
HTTP server Applications making repeated or interactive requests The process stays running and serves JSON endpoints on the configured host and port.
File-based decision Local scripts, scheduled batches, and one-off evaluation A request file is processed without starting a persistent server.
HTTP server through Localtonet Authorized remote applications that cannot connect to loopback directly Our client reaches the local HTTP target and provides a public tunnel address while the tunnel and selected device remain connected.

Connect the working jevos API with Localtonet

Remote requests pass through Localtonet to the jevos HTTP service on a private host.
Localtonet routes remote requests to the working jevos service on the host.

After local health and inference tests succeed, Localtonet can expose the HTTP service to an authorized remote application. Our client establishes an outbound connection to a Localtonet relay, so this workflow does not require inbound router port forwarding, firewall changes, VPN setup, or a public IP address.

Run the Localtonet client on the jevos machine or another device that can reach the configured local target. The simplest arrangement is to run both on the same machine and target 127.0.0.1:8017. If the client runs elsewhere, 127.0.0.1 would refer to that other device, so the target must instead be an address genuinely reachable from the client. Do not change the jevos binding solely by assumption. Plan the network path and access controls first.

1

Install and run the Localtonet client

Install our client on the device that can reach the working jevos HTTP service. Keep jevos running and confirm from that device that the intended local target responds.

2

Authenticate or select the client device

Use the device-specific authentication token associated with the client. Treat the token as a secret and never place it in source code, screenshots, examples, logs, or a public article.

3

Select an available relay server

Choose an available server or region from the current dashboard. Availability can vary, so obtain the current value from the product rather than copying a hardcoded server code from an old guide.

4

Create an HTTP tunnel to jevos

Select an HTTP tunnel and enter the local IP address and port that the client uses to reach jevos. For a client running on the same device with the documented defaults, the target is 127.0.0.1 on port 8017. HTTP process types can use a random subdomain, a supported custom subdomain, or a custom domain, and all serve the target at a public HTTPS address.

5

Start the tunnel

Creating a tunnel does not start it. Use the Start button, then use the assigned public URL shown by the dashboard. The tunnel is available only while the selected client device is connected and the tunnel is running.

6

Test, then stop or delete when finished

Test the health endpoint and one controlled inference request through the assigned URL. Stop the tunnel when temporary access is no longer required, or delete it when the configuration should not be retained.

The complete dashboard sequence and current options are available in our HTTP tunnel documentation. Exact custom-domain DNS instructions are not reproduced here because they should be checked against the current documentation before changing DNS records.

Translate local API URLs to the public tunnel URL

Keep the endpoint paths unchanged and replace only the local origin. If the dashboard assigns a URL represented here as https://your-assigned-host, the corresponding remote endpoints are:

  • https://your-assigned-host/health
  • https://your-assigned-host/v1/models
  • https://your-assigned-host/v1/systemone
  • https://your-assigned-host/dino

The hostname above is intentionally a placeholder, not a real endpoint. Always copy the actual assigned public URL from the dashboard. Do not publish private request data or authentication credentials while testing.

A public URL changes the service's risk profile

A loopback-only API is reachable only from its host. An active HTTP tunnel gives the assigned URL a path to that service. Because the supplied jevos documentation does not establish built-in authentication or authorization, do not expose sensitive decision data or grant untrusted users unrestricted access. Add an appropriate access-control layer and apply least privilege before production use.

Security and production-readiness checklist

Local execution can keep model inference on hardware you control, but remote access still requires a complete security design. The request state may contain customer messages, operational records, or policy information. Protect the endpoint and minimize what each caller can submit.

๐Ÿ” Authenticate callers The supplied jevos evidence does not document built-in API authentication. Place a verified authentication layer in front of the service before allowing untrusted access.
๐Ÿ‘ค Apply least privilege Give only approved applications access and avoid sharing one unrestricted endpoint across unrelated users or environments.
๐Ÿงพ Minimize request data Send only fields needed for the decision. Remove credentials, secrets, unrelated personal data, and unnecessary document content.
๐Ÿšฆ Control workload Define request-size, concurrency, and rate controls in your surrounding application architecture because the supplied evidence does not document production limits.
๐Ÿงช Validate decisions Test representative inputs, monitor error rates, and require human review for uncertain or high-impact outcomes.
โน๏ธ Limit exposure time Stop temporary tunnels after testing. A tunnel remains available only while it is running and its selected client is connected.

Do not place the Localtonet device token in a jevos request, client-side application, repository, or shared configuration file. The token identifies the client device and must remain secret. Likewise, avoid treating an unguessable public URL as authentication.

Separate development and production data. Test remote connectivity first with harmless sample text, verify that only expected paths and methods are used, and observe the behavior when jevos is stopped. A remote application should fail safely if the inference service is unavailable rather than silently approving an action.

The supplied evidence does not establish request logging behavior, audit retention, compliance certifications, uptime guarantees, backup requirements, or a hardened production topology for jevos. Evaluate those requirements explicitly for your organization instead of assuming that a successful local demonstration satisfies them.

Troubleshooting common setup and access problems

uv sync does not run

Confirm that uv is installed and available in the current shell, then confirm that the terminal is in the project directory. The official evidence does not establish a universal uv installation command or operating-system support matrix, so use the installation method appropriate to your platform rather than pasting an unverified command.

The runtime download fails

Confirm that the machine has outbound network access and that the project environment synchronized successfully. Retry from the same repository environment. The runtime is described as machine-specific, so avoid copying an arbitrary runtime from another architecture or operating system unless the project explicitly documents that combination.

The server cannot find the GGUF model

The command uses jevos-q4_k_m.gguf as a relative path. That path resolves from the current working directory. Confirm the filename, download completion, and terminal location. If the file is stored elsewhere, supply its real path to --gguf.

Port 8017 is unavailable

Another process may already be listening on the documented default port. Stop the conflicting process or use the documented --port option to choose an available port. If you change it, update every local request and the Localtonet tunnel target.

The health endpoint refuses the connection

Confirm that the jev serve process is still running and that you are testing from the correct machine. Verify the host and port shown by your server configuration. If the process exited during model loading, inspect its terminal output before investigating Localtonet.

The API returns HTTP 422

Check the request schema and question type. jevos supports yes-or-no questions using "type": "noul". The documented server refuses choice and score questions with status 422. Also confirm that the JSON is valid and required fields are present.

The model ignores a decision rule

Put the rule directly in the question's instructions. Do not rely on the optional criteria field because the implementation accepts but does not read it. Include relevant facts in state and avoid referring to policies that are not present in the request.

Local requests work but the public URL does not

Check the layers in order. Confirm jevos still answers at its local target. Confirm the Localtonet client is connected. Confirm the tunnel has been started, because creation alone does not run it. Finally, confirm that the tunnel points to the same IP address and port used by the successful local test.

The tunnel client runs on another device

On that device, 127.0.0.1 refers to the Localtonet client device itself, not the jevos host. The client needs a network-reachable address for jevos. Changing jevos from loopback to a broader bind can expose it to the local network, so evaluate firewall rules, authorization, and network policy before making that change.

Remote inference is slower than local inference

Compare the response's Server-Timing inference value with end-to-end client timing. If inference time is stable but total duration rises, examine network transit and surrounding application behavior. If inference time itself rises, inspect CPU contention, the chosen thread count, and concurrent workloads.

The output file is not updated

The documented --output behavior never overwrites an existing file. Use a new output filename or intentionally remove or archive the existing file after checking its contents.

Frequently asked questions

What does jevos return for a yes-or-no question?

It returns a noul probability between 0 and 1 for each named question. Higher values indicate stronger support for yes. Your application decides how to interpret that probability and which threshold or review policy to apply.

Can jevos answer multiple-choice or scoring questions?

Not through the documented implementation covered here. It supports yes-or-no noul questions. Choice and score question types are refused with HTTP status 422.

Does jevos require the HTTP server for every decision?

No. The uv run jev decide command can process a request JSON file without starting the persistent server. The HTTP mode is more suitable when applications need repeated or interactive requests.

Can one request contain several questions?

Yes. Add several named entries under questions. They share the same state, which is read once. Each answer appears under the corresponding question name in the response.

Is Docker required or officially documented for this setup?

Docker is not required by the documented quickstart, and the supplied evidence does not establish an official Docker installation. This guide uses the supported sequence of uv sync, runtime download, and jev serve.

Do I need to bind jevos to all network interfaces for Localtonet?

No when our client runs on the same machine. It can target the documented loopback address 127.0.0.1:8017. If the client runs on another device, it needs an address that is reachable from that device, and broadening the bind requires an additional security review.

Does creating a Localtonet tunnel make it active immediately?

No. After creating the tunnel, use the Start button. The assigned address remains available only while the selected client device is connected and the tunnel is running.

Does Localtonet remove the need for API authentication?

No. Tunneling provides network reachability to the local target. It does not make application authorization unnecessary. Because the supplied jevos documentation does not establish built-in authentication, add a suitable access-control layer before exposing sensitive or production workloads.

Can I use a custom domain for the HTTP tunnel?

HTTP tunnels support random subdomains, supported custom subdomains, and custom domains. Consult the current Localtonet documentation before configuring a custom domain because exact DNS requirements should not be copied from an outdated example.

Connect your local jevos API with Localtonet

Verify jevos on 127.0.0.1:8017, secure the application for its intended audience, then create an HTTP tunnel so authorized remote applications can reach the working inference endpoint without inbound router port forwarding.

Get Started Free โ†’

Localtonet is a secure multi-protocol tunneling and proxy platform designed to expose localhost, devices, private services, and AI agents to the public internet supporting HTTP/HTTPS tunnels, TCP/UDP forwarding, mobile proxy infrastructure, file server publishing, latency-optimized game connectivity, and developer-ready AI agent endpoint exposure from a single unified control plane.

support