
Publish one controlled AI gateway while keeping every inference backend private
A self-hosted LLM router gives applications one API while directing requests to different local or network-reachable model backends. The safest public-access pattern is to expose that router, not every Ollama instance, LM Studio server, GPU worker, or provider adapter behind it. This guide explains the architecture, local setup requirements, routing metadata, authentication, verification, failure handling, observability, and Localtonet HTTP tunnel workflow. Localtonet provides the public connection to the gateway, while the router remains responsible for model selection and request policy.
๐ What's in this guide
The correct network boundary for a self-hosted LLM router
A multi-model deployment often starts as several independent services. One machine may run Ollama, another may run an application around LM Studio, and a GPU server may host a specialized inference engine. An organization may also have adapters for remote model providers. Connecting every application directly to every backend creates a difficult network and security problem. Each client needs backend addresses, credentials, model identifiers, retry behavior, and knowledge of which service can handle a particular request.
An LLM router or reverse proxy consolidates that complexity behind one HTTP API. Clients send requests to the router. The router validates the request, evaluates whatever routing policy has been configured, chooses an eligible backend, forwards the request, and returns the response. Depending on the router software, its decision might use an explicitly requested model, a task label, estimated complexity, cost policy, backend health, capacity, or a fixed rule. These capabilities are properties of the selected router and its configuration, not properties of the network tunnel.
With Localtonet, the intended boundary is straightforward: the Localtonet client reaches the router's local HTTP listener, and the HTTP tunnel publishes that single listener. Backend model services remain on loopback addresses, private container networks, or protected LAN addresses that are reachable by the router but are not assigned their own public tunnels.
Remote application
|
| HTTPS request
v
Localtonet public address
|
| HTTP tunnel to the selected local target
v
Authenticated LLM router
|
+-- Private backend: Ollama
+-- Private backend: LM Studio
+-- Private backend: GPU worker
+-- Private backend: another approved model service
This architecture reduces the publicly reachable surface. It also gives us a practical control point for authentication, request validation, routing policy, rate controls, audit metadata, and error normalization. It does not make the private services inherently secure, however. Backend listeners should still be bound and filtered appropriately because other processes or devices on the same host or LAN may be able to reach them.
What Localtonet does and does not do
Localtonet transports HTTP traffic between a public address and the local IP address and port selected for the tunnel. It does not inspect prompts to decide which model should receive them, calculate model cost, manage an inference queue, translate incompatible model APIs, or determine whether a backend has enough GPU memory. Those responsibilities belong to the LLM router and the inference stack.
This separation is useful. The router can evolve independently from the connectivity layer. You can modify routing rules, add a model worker, remove an unhealthy backend, or change an alias without giving clients a new network address, provided the gateway's public API contract remains stable.
A model server may have been designed for trusted local use and may not enforce authentication, authorization, payload limits, or tenant separation. Publishing each backend independently also allows clients to bypass gateway policy. Expose only the authenticated router unless a separate backend has a documented, reviewed reason to be public.
Prerequisites and decisions to make first
This guide is router-agnostic because no specific open-source router was selected for the workflow. Installation commands, configuration file names, environment variables, default ports, health paths, and supported operating systems differ between router projects. We therefore do not invent a universal installation command. Install your chosen router using its maintained package, container, or source procedure, then apply the architecture and verification sequence below.
Before adding a tunnel, the complete inference path must work locally. You need a running router, at least one backend the router can reach, a documented API path for a harmless test request, and an authentication mechanism suitable for every client that will use the public endpoint.
| Requirement | What to establish | Why it matters |
|---|---|---|
| LLM router | A supported installation that listens on a known local IP address and port | This is the only service the HTTP tunnel should target. |
| Model backend | At least one locally reachable and tested inference endpoint | The gateway cannot complete requests if all upstream backends are unavailable. |
| API contract | The exact route, request schema, streaming behavior, and response schema clients will use | An OpenAI-compatible label does not guarantee that every optional field or endpoint behaves identically. |
| Gateway authentication | A documented token, key, identity proxy, or other supported access-control mechanism | The public URL should not become unauthenticated access to compute resources or model data. |
| Backend restrictions | Loopback binding, a private subnet, container-network isolation, and appropriate firewall policy | Clients should be unable to bypass the router and call model workers directly. |
| Localtonet client | The client installed on a device that can reach the router's listener | The tunnel's local target is resolved from the client device's network perspective. |
| Operational limits | Request size, concurrency, timeout, queue, and model context policies supported by the chosen stack | Inference requests can consume substantial memory, GPU time, and connection duration. |
Choose the Localtonet client location carefully
The Localtonet client can run on the same machine as the router or on another device that can reach it. When both run on the same host, a loopback target is often the smallest network boundary. If the client runs elsewhere, the router must listen on an interface reachable from that client, and host firewall rules should allow only the necessary source network or device.
Container deployments require particular attention. A loopback address inside one container refers to that container, not automatically to the host or another container. The router, model backend, and Localtonet client must have an intentional network path between them. Use the addressing mechanism documented by your container platform rather than assuming that a host loopback address will work from every container.
Define a stable client-facing model contract
Clients should ideally request stable aliases such as a general-purpose model, a coding model, or an embedding model rather than internal worker names. The router can map an alias to one or more eligible backends. This prevents internal hostnames, vendor-specific identifiers, and GPU topology from becoming part of the public API contract.
Decide which request metadata clients may control. A client-selected model name, maximum output length, tool definition, response format, or routing hint can materially affect cost and workload. The router should reject unsupported values rather than passing arbitrary instructions to every backend.
Configure the self-hosted router as the only gateway

The exact syntax depends on the project, but the configuration goals are consistent. Treat the following as an implementation checklist, then express each item using settings that your router officially supports.
Install the router using its maintained installation path
Use the router project's documented package, container image, or source-build process. Confirm its supported runtime and platform requirements. Do not copy commands intended for another router merely because both advertise an OpenAI-compatible interface.
Register only private backend addresses
Configure the router with the loopback, private LAN, or private container-network address of each approved model service. Keep backend credentials out of source control and client-visible responses. Test reachability from the router's own process or container.
Create stable model aliases and routing policy
Map client-facing names to eligible backends. If the router supports cost-aware or complexity-aware routing, define the candidate pool, compatibility rules, quality threshold, health criteria, and fallback behavior explicitly. A lower advertised token price does not necessarily mean a lower total task cost if a model uses more tokens or requires retries.
Require authentication at the gateway
Enable the router's supported authentication mechanism or place a reviewed authentication layer directly in front of it. Use separate credentials where the stack supports them, limit their permissions, rotate them safely, and never embed real secrets in examples, logs, browser code, or public repositories.
Set resource and request controls
Configure only controls supported by the selected router, such as request-size limits, allowed models, output limits, concurrency, queues, deadlines, and per-client restrictions. Align upstream timeouts with realistic model startup and generation times so one layer does not abandon work while another continues consuming resources.
Start the router and inspect startup status
Verify that the process remains running, listens on the intended address, loads the expected model mappings, and can connect to at least one backend. Resolve configuration and backend errors before creating public access.
Design routing metadata deliberately
Useful request metadata can include a client identity, application name, request ID, model alias, workload class, or tenant identifier. Only accept metadata that the router understands and validates. Do not allow an untrusted client to submit arbitrary internal backend URLs, raw provider credentials, unrestricted administrative flags, or a forced route that bypasses policy.
The gateway should generate a request ID when the client does not provide an acceptable one. Include that identifier in gateway logs and return it in an appropriate response field or header if the selected router supports doing so. This lets an operator trace a remote failure through the gateway without logging an entire sensitive prompt.
Make failure behavior predictable
Fallback is not simply โtry every model.โ A request might use tools, structured output, images, embeddings, a large context, or a streaming mode that only some backends support. Build an eligible pool first, then select or retry within that pool. If no compatible backend is healthy, return a controlled error rather than silently downgrading to an incompatible model.
Automatic retries also need boundaries. Retrying a non-idempotent agent action can execute a tool twice. Retrying after partial streaming output can produce duplicate or conflicting text. Prefer retries before response bytes are committed, and distinguish connection failures from valid model rejections. The exact controls depend on the router, but the policy should be explicit and tested.
Different gateways and inference servers may implement different endpoint subsets, fields, streaming formats, tool-calling behavior, model-list responses, error bodies, and token accounting. Test the exact operations your client uses. Do not assume compatibility based only on the name of the API style.
Verify the complete API locally before tunneling it

Local verification separates router problems from tunnel problems. If the same request does not work against the router's local address, adding a public URL will not repair the backend configuration, model name, authentication, or request schema.
Set a local base URL using the actual listener shown by your router. The placeholder below is intentionally not a claimed default port:
export ROUTER_BASE_URL="http://127.0.0.1:<router-port>"
Do not paste that placeholder unchanged. Replace it with the address and port from your router configuration. If the Localtonet client will run on another machine, repeat the test from that machine using the router's permitted private address. A successful test from the router host alone does not prove that a separate tunnel client can reach it.
Test liveness separately from inference
If the router documents a health or readiness endpoint, call that exact route and verify its documented status. We cannot provide one universal path because router projects differ. A useful health check should establish whether the gateway process is ready to accept traffic, while a deeper readiness check may also consider backend availability.
Next, send a minimal authenticated inference request through the same API route that remote clients will use. For an OpenAI-compatible router, a typical chat-style payload has the following shape, but the endpoint, fields, and model alias must be confirmed against the chosen router:
curl --fail-with-body \
--request POST \
"${ROUTER_BASE_URL}/v1/chat/completions" \
--header "Authorization: Bearer ${ROUTER_API_TOKEN}" \
--header "Content-Type: application/json" \
--data '{
"model": "approved-model-alias",
"messages": [
{
"role": "user",
"content": "Reply with the word ready."
}
],
"stream": false
}'
Keep the real token in a protected secret mechanism appropriate for your environment. Do not store it in shell history, commit it to a repository, or place it in frontend JavaScript. If your router uses a different authentication header or payload, use its documented form instead.
Run negative tests as well as successful tests
A successful response proves only one path. Before public exposure, test that the gateway rejects a missing credential, an invalid credential, an unknown model alias, an unsupported request field where strict validation is expected, and a request that exceeds a configured limit. Temporarily make one nonessential backend unavailable and confirm that the observed behavior matches your fallback policy.
| Local test | Expected result | Problem it can reveal |
|---|---|---|
| Documented readiness check | Router reports ready according to its documentation | Startup failure, invalid configuration, or unavailable required dependency |
| Valid authenticated inference | A response arrives through the router from an eligible backend | Incorrect route, schema, model alias, credentials, or backend connectivity |
| Missing authentication | Request is rejected before inference begins | Accidentally unprotected gateway |
| Unknown model alias | Controlled client error without internal topology disclosure | Unsafe pass-through behavior or weak input validation |
| Backend outage | Documented fallback or controlled unavailable response | Retry storms, incompatible fallback, or misleading success status |
| Streaming request, if supported | Chunks arrive and terminate in the client-expected format | Buffering, timeout, disconnect, or protocol compatibility problems |
Expose the verified router with a Localtonet HTTP tunnel

Once the gateway works from the Localtonet client device, create an HTTP tunnel that points only to the router. Our client establishes an outbound connection to a Localtonet relay server. You do not need inbound router port forwarding, a public IP address, firewall changes for unsolicited internet traffic, or a separate VPN setup for this workflow.
HTTP tunnels support a Random Sub Domain, Custom Sub Domain, or Custom Domain process type. These process types serve the same local content at a public HTTPS address. Availability can vary, so use the choices shown for your account and current dashboard. Exact custom-domain DNS instructions should be taken from the current Localtonet documentation rather than guessed.
Install and run the Localtonet client
Run the client on the router host or on another device that can reach the router's private listener. Keep the client active for as long as remote applications need the gateway.
Authenticate or select the client device
Use the device-specific auth token supplied through our platform and select the intended device. Treat that token as a secret. Never publish it in configuration examples, screenshots, logs, or client applications.
Select an available relay server or region
Choose from the values currently available in the product or dashboard. Do not hardcode a server code copied from another account or an old article because available values can vary.
Create an HTTP tunnel to the router
Select the appropriate HTTP process type and set the local target to the router's reachable IP address and port. Target the gateway listener, not an Ollama, LM Studio, GPU-worker, or other backend listener.
Start the tunnel
Creating a tunnel does not mean it is running. Use the Start button and confirm that the selected Localtonet client is connected.
Use and verify the assigned public address
Replace the local base URL in the same authenticated test with the assigned public HTTPS URL. Confirm successful inference, authentication rejection, expected routing, and streaming behavior where applicable. Stop or delete the tunnel when public access is no longer required.
The current workflow and available controls should be checked in the
Localtonet HTTP tunnel documentation.
The public address exists as a path to your local gateway, but the gateway's own API paths remain unchanged. If the local request uses a documented path such as /v1/chat/completions, the remote client normally appends that same path to the assigned public base URL.
Keep gateway authentication enabled and test that unauthenticated requests fail. The tunnel publishes connectivity to the selected service. Your router or its adjacent authentication layer must decide who may use models, which models they may request, and what limits apply.
Verify that the backends are still private
After the tunnel is running, inspect the Localtonet configuration and confirm that only the router's listener is selected. Search application configuration and documentation for any backend address accidentally given to remote clients. From an external network, only the assigned gateway address should be part of the supported access path.
Do not rely solely on the absence of another Localtonet tunnel. A model service bound to all network interfaces could still be reachable from the local LAN or through unrelated firewall and cloud rules. Review listener bindings, host firewall policy, container port publishing, cloud security groups where applicable, and router administration interfaces.
Security controls for a public AI gateway
LLM endpoints deserve the same defensive treatment as other compute-intensive authenticated APIs, with additional attention to prompt content, model output, tool execution, and potentially expensive workloads. The network tunnel should be one layer in a broader design.
Authenticate before allocating inference work
Reject invalid credentials as early as possible. Authentication should occur before a request enters a model queue, triggers model loading, or consumes a remote provider quota. If the router supports per-client credentials or scopes, assign the minimum models and actions each client needs. A monitoring probe should not require administrative access, and a low-trust application should not automatically receive every model alias.
Keep credentials separated by purpose
The Localtonet device auth token identifies the client device that runs the tunnel. It is not an LLM API credential and should never be distributed to API consumers. Gateway credentials authorize use of the AI API. Backend credentials allow the router to reach protected model services or providers. Keep these three roles separate so rotating one does not unnecessarily expose or disrupt the others.
Constrain requests, not just request rates
A small number of requests can still consume significant resources if they include large contexts, request long outputs, invoke tools, or trigger expensive models. Where supported by the chosen router, apply limits to payload size, context, output, concurrency, model access, and queue depth. Use server-side policy rather than trusting clients to submit conservative values.
Protect prompt and response data
Prompts may contain source code, personal data, internal documents, credentials, or tool results. Avoid logging complete bodies by default. Prefer operational fields such as timestamp, request ID, authenticated client, requested alias, selected backend category, latency, status, token counts when reliably reported, and error class. If prompt capture is necessary for a specific debugging event, restrict access and establish a deletion process.
Separate inference from administration
If the router has an administrative dashboard, configuration API, metrics listener, profiling endpoint, or backend registration endpoint, do not assume it should share the public inference tunnel. Bind administrative interfaces privately and expose only the client-facing API required for inference. An unrestricted configuration endpoint could let an attacker alter routes, extract secrets, or point the gateway at an unintended destination.
Treat tools as a separate trust boundary
A model gateway that can invoke tools is more than a text-generation endpoint. Tool calls may read files, query databases, send messages, or modify external systems. Authenticate tool-capable clients, allowlist tools, validate arguments, isolate execution, and require confirmation for sensitive actions. A tunnel should never be described as bypassing security policy or authorization.
Observability, failure handling, and routine operations
A gateway hides backend topology from clients, but operators still need enough visibility to understand routing quality and system health. Monitor the complete request path rather than treating an HTTP success from the router process as proof that inference is healthy.
Measure each stage of the request
Useful measurements include authentication rejection counts, queue delay, routing-decision time, time to first response byte for streaming, total duration, selected model alias, backend outcome, retry count, fallback count, and cancellation status. Token usage can be valuable when the backend reports it consistently, but do not assume every implementation measures tokens in the same way.
Distinguish gateway errors from backend errors. A malformed request, unauthorized credential, unavailable model, backend timeout, capacity rejection, and client disconnect require different responses. Returning the same generic server error for all of them makes automated retry behavior dangerous and troubleshooting unnecessarily difficult.
Use correlation without exposing internal topology
Assign a request ID at the gateway and propagate it to backend requests where supported. Client-facing errors can include that request ID without revealing private hostnames, container names, credentials, stack traces, or internal network addresses. Operators can then locate the relevant gateway and backend events.
Plan for cold starts and long responses
A local model may need time to load into memory. Generation duration also varies with prompt length, requested output, hardware, model size, and contention. Set client, gateway, and backend deadlines intentionally. A client timeout that is shorter than model startup can appear as a tunnel failure even when the backend eventually completes the request.
For streaming APIs, test how cancellation propagates when the remote client disconnects. Ideally, abandoned work should stop where the stack supports cancellation. Otherwise, repeated disconnects can leave the GPU processing responses nobody will receive.
Maintain the tunnel lifecycle
The tunnel is available only while the selected Localtonet device is connected and the tunnel is running. After maintenance or a host restart, confirm the Localtonet client, LLM router, and required backends are all active. Creating a tunnel in the dashboard alone does not start it.
When remote access is temporary, stop the tunnel after the work is complete. Delete obsolete tunnels and retire unused gateway credentials. If a device auth token or gateway secret may have been disclosed, rotate or replace it using the relevant supported workflow rather than attempting to hide it only in logs.
Validate changes before broad use
A model update can alter output format, context support, tool behavior, resource consumption, and latency. A routing-rule update can send requests to an unexpected backend even though the public address remains healthy. Test configuration changes locally, exercise both success and failure paths, then repeat a controlled request through the public URL.
Cost-aware routing also needs outcome monitoring. Selecting the model with the lowest nominal price can be counterproductive if it produces longer outputs, requires repeated attempts, or fails tasks that a more capable model completes once. Evaluate total request trajectories and successful outcomes, not only the listed price of a single model call.
Troubleshooting the gateway and tunnel path

The public URL is unreachable
Confirm that the Localtonet client is connected and that the tunnel was started. A configured but stopped tunnel is not publicly available. Verify that the selected device is the one running the client and that it remains online. Check the current tunnel state in our dashboard before changing the LLM router.
The tunnel returns a connection or upstream error
Test the router from the Localtonet client device using the exact local target address and port. If that test fails, correct the listener, firewall, container network, or service state first. A common container mistake is targeting a loopback address that belongs to the Localtonet client container rather than the router container or host.
Local requests work, but remote requests are unauthorized
Compare the remote request with the known-good local request. Confirm that the authorization header reaches the gateway, that the credential is valid for the requested model, and that the client is not accidentally sending a Localtonet device token as an LLM API token. Do not disable authentication to make the test pass.
Remote requests work without credentials
Stop the tunnel while investigating. Confirm that authentication is enabled on the exact listener targeted by the tunnel. Some applications expose separate development and production listeners, or protect one route while leaving another open. Test missing and invalid credentials locally before restarting public access.
The router reports an unknown model
Use the client-facing alias configured in the gateway, not necessarily the backend's native model name. Check for case differences and environment-specific mappings. Avoid exposing a raw backend model inventory unless that behavior is intentional and authorized.
Requests time out only for larger models
Measure queue time, cold-start time, time to first output, and total generation time. Review the documented timeout settings in the client, router, and backend. Also check memory pressure, model loading, context size, and concurrency. Do not assume that every timeout originates in the tunnel.
Streaming works locally but not in the application
First repeat the streaming call directly against the public URL with a simple HTTP client. Confirm that the router emits the documented streaming format and terminates it correctly. Then inspect application-side buffering, response parsing, cancellation, and timeout behavior. Compatibility varies across routers, SDKs, and inference servers, so test the exact combination you deploy.
The router selects an unexpected backend
Record the requested alias, validated routing metadata, eligible candidate set, health state, and final selection without logging sensitive prompt bodies. Check whether the request required a capability unavailable on the expected backend. If the router uses complexity-aware or cost-aware selection, review its documented decision inputs and policy rather than attributing the choice to Localtonet. The tunnel carries the request but does not choose the model.
Failures continue after the model service recovers
Determine how the router marks a backend unhealthy and how it returns that backend to service. Recovery checks, retry intervals, and circuit-breaking behavior are router-specific. Confirm readiness directly from the router's network location, then use the router's documented recovery procedure. Avoid repeatedly restarting the public tunnel when the actual unhealthy state is inside the routing layer.
Frequently asked questions
Does Localtonet choose which LLM handles a request?
No. Localtonet provides the public-to-local HTTP connection. The self-hosted router chooses the backend according to its own model aliases, health checks, compatibility rules, cost policy, task classification, or other supported configuration.
Should Ollama or LM Studio receive a separate public tunnel?
Not for this architecture. Keep model backends on loopback, a private LAN, or a private container network and expose only the authenticated gateway. A separate public backend should require a specific reviewed use case and appropriate application-level controls.
Does the Localtonet public URL replace gateway authentication?
No. The public address provides connectivity to the selected local service. The router or an adjacent authentication layer must validate API consumers and enforce model access, request limits, and other authorization policy.
Can the Localtonet client run on a different machine from the router?
Yes, provided the client device can reach the router's local IP address and port. Restrict that private listener with appropriate firewall and network policy. Test the router from the client device before creating the tunnel.
Which port should the self-hosted LLM router use?
Use the listener port documented or explicitly configured for your chosen router. There is no universal port for all LLM gateways, so this guide intentionally does not invent one. Configure the Localtonet HTTP tunnel with that verified local target.
Can clients continue using an OpenAI-compatible SDK?
They can if the selected router implements the endpoints and behaviors that the SDK actually uses and the SDK allows its base URL to be changed. Verify authentication, request fields, streaming, tool calls, structured output, model listing, and error handling because compatibility differs between implementations.
Does creating the tunnel immediately make the API available?
No. The tunnel must be started, and the selected Localtonet client must be connected. The router and its required model backends must also be running. The endpoint stops being available when the tunnel is stopped or the client device disconnects.
Should prompts and responses be logged for troubleshooting?
Not by default. Prefer request IDs, client identity, model alias, selected backend category, timing, token metadata where reliable, and error classes. Prompt or response capture should be a deliberate, access-controlled debugging action with an appropriate retention and deletion policy.
Publish one controlled LLM gateway with Localtonet
Verify your authenticated router locally, keep its model backends private, and then connect the single gateway listener to a Localtonet HTTP tunnel for remote HTTPS access.
Get Started Free โ