
Find where capable AI models are being used for work that may not require them
Request totals and token counts can show that AI usage is growing, but they do not explain whether the growth comes from valuable complex work, inefficient agents, broad default-model selection, or repeated attempts to finish simple tasks. In this guide, we build an observability model for a self-hosted AI gateway that connects each request to a user, application, task, conversation, turn, model, cost, and outcome. We then show how to investigate possible model overuse without treating a heuristic as proof. Finally, we explain how a Localtonet HTTP tunnel can make the gateway dashboard remotely reachable while keeping application-layer AI analytics separate from tunnel connectivity.
📋 What's in this guide
Why AI model usage needs more context
AI gateway observability should answer more than how many requests crossed the gateway. It should explain who or what initiated those requests, what work was being attempted, which model handled it, how much interaction was required, and whether the result met an operationally meaningful standard.
Aggregate request and token charts remain useful. They can reveal traffic growth, usage spikes, unusual output volume, and changing cost. However, the same totals can represent very different situations. A sustained research workflow may legitimately use a capable reasoning model and several turns. A formatting operation may send a few hundred requests to that same model only because it is configured as the application default. An agent may also turn one user action into many hidden model calls, making per-user request counts misleading unless the gateway records the relationship between the user action, conversation, agent run, and individual calls.
“Model overuse” should therefore be treated as an investigation category, not a final verdict. A large or highly capable model is not automatically the wrong choice for a short prompt, and a less expensive model is not automatically sufficient. Requirements such as accuracy, structured-output reliability, tool selection, language coverage, context length, safety behavior, and response consistency can justify a model that appears excessive from token counts alone.
A model-fit rule should identify workloads worth examining. It should not silently replace models before the team compares output quality, latency, task completion, retries, total conversation cost, and any application-specific acceptance criteria.
The practical objective is to find repeatable patterns. For example, a team might discover that one application routes all summarization and formatting requests to its most capable reasoning model. Most of those conversations finish in one turn and have predictable output requirements. That combination creates a sensible candidate for controlled evaluation with another model. By contrast, a multi-turn debugging conversation with tool calls, corrections, and a verified final patch may be operating exactly as intended.
Separate AI telemetry from tunnel connectivity
A clean design separates the data plane that processes AI requests, the observability pipeline that generates usage insights, and the connectivity layer that provides approved remote access. Each layer has a different responsibility and should be monitored independently.
| Layer | Primary responsibility | Typical data | What it should not be assumed to do |
|---|---|---|---|
| AI application or agent | Initiates work and supplies business context | User reference, task reference, conversation ID, agent-run ID, requested model | It should not be assumed to produce complete gateway-wide analytics unless it is explicitly instrumented to do so. |
| Self-hosted AI gateway | Accepts model requests, applies routing policy, calls providers or model servers, and records request results | Selected model, token usage, latency, status, routing decision, retry and error information | It cannot reliably infer every business outcome or user identity when the application does not provide that context. |
| Telemetry pipeline | Normalizes, stores, aggregates, and analyzes gateway events | Task groups, per-user usage, conversation totals, turn distributions, model-fit candidates | It should not be treated as an infallible judge of task complexity or output quality. |
| Analytics dashboard | Presents trends and supports investigation | Filters, distributions, drill-downs, candidate overuse views, quality and cost comparisons | A dashboard does not replace controlled evaluation or authorization controls. |
| Localtonet HTTP tunnel | Makes a local HTTP service reachable through a public HTTPS address | Tunnel lifecycle, local target, assigned public address, device connectivity | Localtonet does not generate semantic task classifications or model-usage insights for the gateway. |
This separation prevents a common observability mistake. A tunnel can provide connectivity to an internal dashboard, but connectivity metadata alone cannot explain whether a prompt was a code review, a summary, or one step in an agent workflow. The application and gateway must create the semantic telemetry. Localtonet can then provide remote reachability for the resulting web interface without requiring inbound router port forwarding, firewall changes, VPN setup, or a public IP address.
The gateway should also avoid making identity decisions from network information alone. Several users can share an address, automated workloads can run behind the same egress path, and one user can access an application from multiple networks. Use an authenticated application identity or a controlled pseudonymous identifier instead of treating an IP address as a user identity.
Define boundaries before collecting data
Decide which component owns each field. The application usually knows the authenticated user, business task, workspace, and conversation. The gateway knows the requested model, selected model, timestamps, provider response, token counts when supplied, and whether routing or retries occurred. A downstream evaluator may know whether the output passed a test, was accepted by a user, or completed a workflow.
When a field is unavailable, store it as unknown rather than manufacturing a value. This is especially important for provider-reported token counts, model pricing, task categories, and quality scores. A fabricated zero can distort totals, while an explicit unknown preserves the limitation.
Design a privacy-conscious telemetry event model
An effective event model supports several levels of analysis without requiring raw prompts in every record. Start with stable identifiers and operational metadata. Add content-derived attributes only when they are necessary, lawful, protected, and proportionate to the intended analysis.
The following JSON is an illustrative schema, not a Localtonet configuration and not a requirement for any specific gateway. Field names and availability must be adapted to the gateway, model provider, identity system, and storage platform in use.
{
"event_id": "generated-event-identifier",
"occurred_at": "timestamp",
"actor": {
"user_ref": "pseudonymous-user-reference",
"team_ref": "team-reference",
"application_ref": "application-reference",
"agent_ref": "optional-agent-reference"
},
"work": {
"task_ref": "task-reference",
"task_category": "summarization",
"classification_method": "application-supplied",
"conversation_ref": "conversation-reference",
"turn_number": 1
},
"routing": {
"requested_model": "model-requested-by-client",
"selected_model": "model-used-for-call",
"route_rule_ref": "optional-routing-rule",
"fallback_used": false
},
"usage": {
"input_tokens": 0,
"output_tokens": 0,
"latency_ms": 0,
"cost_amount": null,
"cost_currency": null
},
"result": {
"status": "completed",
"retry_count": 0,
"quality_signal": null,
"validation_passed": null
},
"privacy": {
"prompt_stored": false,
"response_stored": false,
"retention_class": "operational-analytics"
}
}
Prompts and responses can contain personal data, source code, customer records, access tokens, internal URLs, secrets, or regulated information. Prefer metadata, pseudonymous references, bounded classifications, and explicit quality signals. If content capture is genuinely required, apply access controls, redaction, retention limits, and organizational review appropriate to the data.
Identity and workload fields
A user_ref should support investigation without exposing a direct identifier to every analyst. One common design is to create a stable pseudonymous reference in the application and keep the re-identification mapping in a more restricted system. Service accounts and agents should have separate identity types so that automated traffic is not incorrectly attributed to a person.
Record both application and team context when available. A user may interact with several AI-enabled tools, and the same gateway may serve production applications, internal experiments, scheduled jobs, and autonomous agents. An application reference makes it possible to determine whether model selection comes from individual behavior or a shared configuration.
Task, conversation, and turn fields
A task category gives meaning to model usage. Categories can include coding, research, writing, summarization, classification, extraction, data analysis, or organization-specific work. These categories are examples, not a universal taxonomy. Keep the taxonomy small enough to interpret, document what each category means, and include an unknown category rather than forcing uncertain requests into a misleading label.
Store how the classification was produced. An application-supplied category may be highly reliable when the endpoint performs one known function. A rule-based category may be adequate for deterministic routes. A model-generated classification can be useful, but it introduces its own cost, latency, privacy, and accuracy considerations. Analysts should be able to separate these methods.
Conversation and task identifiers are related but not always identical. A conversation may contain several tasks, while an agent run may issue several model requests for a single task. Define these relationships explicitly. The turn number should describe its scope, such as a user-visible exchange or an internal agent step, so that a “five-turn task” has a consistent meaning.
Routing, usage, and outcome fields
Record both the requested and selected model. If only the selected model is stored, investigators cannot tell whether an application requested it directly, the gateway chose it through a policy, or a fallback changed the destination. A route-rule reference can make broad defaults visible without copying the entire routing configuration into each event.
Token values should preserve their provenance. Some model servers or providers report usage directly, while other integrations may not. Do not silently estimate missing values and mix them with provider-reported counts unless the data model labels the estimation method.
Cost should be derived from a versioned pricing table appropriate to the provider, model, time, and contract. Pricing is not static, and different organizations can have different commercial terms. When reliable pricing is unavailable, retain token and request metrics and mark cost as unknown.
Outcome is the most important field and often the least complete. Useful signals can include a schema validation result, unit-test result, task completion status, accepted edit, user rating, escalation, regeneration, or rollback. No single signal works for every task. A summary may need human review, while generated code can sometimes be tested automatically.
Instrument the self-hosted AI gateway

Instrumenting a gateway is not merely enabling access logs. The gateway must preserve correlation across application requests, routing decisions, provider calls, retries, and outcomes. Because gateway implementations differ, there is no universal command, port, file path, or environment variable that can be stated safely. Use the configuration mechanism supported by the specific gateway and apply the following implementation sequence.
Define the questions and decision criteria
Write down the investigations the telemetry must support. Examples include identifying high-capability model usage by task, finding agents with unusually high calls per task, comparing one-turn summarization workloads across models, and verifying whether a routing change preserves quality. Define what evidence is required before changing a route.
Create correlation identifiers at the application boundary
Generate or propagate references for the user or service, application, task, conversation, and request. Do not place credentials or direct personal information in those identifiers. Ensure automated agents have distinct references and that identifiers remain consistent across retries and provider calls belonging to the same task.
Capture the routing decision
Record the requested model, selected model, relevant routing-rule reference, fallback state, and provider or model-server destination when that information is available and approved for telemetry. This establishes why a request reached a particular model.
Record request and response measurements
Capture timestamps, latency, status, retry count, and provider-reported input and output usage where supported. Preserve missing values as unknown. If streaming is used, define whether latency means time to first output, total completion time, or both.
Attach task and outcome context
Add an application-supplied task category where possible. Associate validation, user feedback, completion, or other quality signals after the response is evaluated. Late-arriving outcomes can be joined to the original event through the task or request reference.
Normalize and validate events
Validate required fields, accepted categories, timestamps, numeric ranges, identifier formats, and schema versions before storage. Quarantine malformed events rather than silently dropping them. Track instrumentation coverage so analysts know which applications and models have complete data.
Aggregate by request, turn, conversation, and task
Create views at multiple scopes. Request-level records support debugging, while conversation and task aggregates reveal the total tokens, latency, retries, calls, and cost required to finish work. Keep the underlying scope explicit in every metric.
Verify with controlled test traffic
Run known test cases through each important application path. Confirm identity propagation, task classification, requested and selected model values, turn ordering, retry handling, token availability, and outcome updates. Avoid using sensitive production prompts as test fixtures.
Handle retries and agent fan-out correctly
Retries can inflate apparent demand if every provider attempt is counted as a separate user request. Preserve both scopes: the logical gateway request and each downstream attempt. Similarly, an agent may use planning, retrieval, tool selection, verification, and final-answer calls. Those calls should remain visible, but they should roll up to a common agent run and task.
This distinction supports two different questions. Operations teams may need downstream attempt counts to diagnose provider errors. Product teams may need logical tasks completed per user. Combining the two into one request metric makes both analyses less reliable.
Track telemetry quality as a first-class metric
Before interpreting overuse charts, measure missing user references, unknown task categories, absent token counts, unjoined outcomes, duplicate events, and out-of-order turns. A dashboard that shows precise percentages over incomplete data can create false confidence. Display coverage beside every important analysis, especially when comparing applications with different instrumentation maturity.
Detect possible model overuse without oversimplifying it
The strongest overuse signal combines task simplicity, model capability, total task cost, turn behavior, latency, and outcome quality. No single metric establishes that a model is unnecessary. Build a candidate view that brings these dimensions together and lets an operator inspect the underlying workflow.
Start with cohorts, not isolated requests
Group comparable work before drawing conclusions. Useful cohorts may share an application, task category, model, route rule, input-size band, output format, and time period. Comparing unrelated tasks can make a model appear inefficient simply because it receives more demanding work.
Within each cohort, examine median and tail behavior rather than relying only on averages. A generally efficient workflow can contain a small group of extreme agent loops, while an average can hide that concentration. Also inspect the number of distinct users, agents, and applications. A single misconfigured agent may account for a large portion of a model’s usage.
| Signal | Possible interpretation | Required follow-up |
|---|---|---|
| Simple task category with a capable model | A broad default route may be applying more capability than routine work appears to need. | Compare quality and reliability against an approved alternative using representative evaluations. |
| Many turns for routine tasks | The prompt, interaction design, model, or agent workflow may be failing to complete work directly. | Inspect retries, corrections, validation failures, and the role of each turn before changing the model. |
| High output tokens for bounded output | Instructions or output controls may be too loose. | Check expected output length, truncation, structured-output validation, and user acceptance. |
| One user or agent dominates usage | A specialized workload, automation loop, or configuration error may be concentrated in one actor. | Review the workload owner, task outcomes, schedules, and agent stopping conditions. |
| Requested model differs from selected model | A routing rule or fallback may be changing the application’s intended choice. | Inspect the route decision, provider health, policy version, and fallback reason. |
| Low latency but poor completion outcomes | A fast or inexpensive route may not be effective if users retry or abandon the result. | Measure the whole task, including follow-up turns, regeneration, and validation. |
Calculate the full cost of a task
The first model call is only one part of task cost. A useful task aggregate includes all related model attempts, input and output usage, waiting time, retries, validation calls, and follow-up turns. If an inexpensive model requires repeated corrections while another model completes the task once, request-level cost can point in the wrong direction.
Conversation length also needs interpretation. Long research and debugging conversations may be productive. The more actionable pattern is a repeated multi-turn trajectory for a narrowly defined task that should have a bounded result. Even then, investigate whether ambiguity, missing context, poor prompts, tool failures, or application design caused the extra turns.
Use a reviewable candidate score
Teams can rank candidates with an internal score, but every component should remain visible. A score might consider whether the task belongs to a known routine category, whether a high-capability model was selected, whether the task usually completes in one turn, whether an approved alternative has passed evaluation, and whether quality signals are available. The exact formula is organization-specific and should not be represented as a universal model-overuse standard.
Avoid making the score a leaderboard for employees. User-level analysis is useful for finding application defaults, training gaps, or automated loops, but it can be misinterpreted when workloads differ. Review results with task and application context, limit access to the data, and document the operational purpose of the analysis.
Change routing through controlled evaluation
Once a candidate is identified, replay representative, appropriately protected test cases against the current model and an approved alternative. Compare application-specific quality, completion rate, latency, total task usage, and failure modes. Include difficult examples and edge cases, not only the most convenient prompts.
Roll out changes gradually where the application supports it. Monitor route selection, quality signals, retries, escalations, and user feedback. Preserve a fallback path for failures that the application can safely detect. A successful routing change is not merely cheaper per request. It should maintain the required outcome while improving the relevant operational objective.
Protect prompts, identities, and analytics endpoints
AI telemetry can become more sensitive than ordinary service metrics because it may reveal what individuals are working on, which repositories or customers are involved, and how internal agents are configured. Minimize data before attempting to secure a larger collection.
- Prefer metadata over content: Store task category, token counts, model, outcome, and pseudonymous references when those fields answer the question.
- Separate identity mapping: Keep any mapping between pseudonymous actor references and direct identities in a more restricted system.
- Redact before storage: If content-derived analysis is required, remove secrets and unnecessary personal or business data before the event reaches general analytics storage.
- Set retention by purpose: Debugging records may not need the same retention as long-term aggregate trends. Define and enforce separate retention classes.
- Restrict raw-event access: Most dashboard users may only need aggregate views. Limit detailed prompt, response, identity, and request records to approved roles.
- Audit configuration changes: Track changes to routing policies, task classifiers, pricing tables, and dashboard permissions so trends can be interpreted correctly.
- Protect exports: CSV files, notebooks, and screenshots can escape dashboard controls. Apply the same handling rules to exports as to the underlying data.
If a dashboard becomes reachable from the public internet, protect it with appropriate application authentication and least-privilege authorization. Do not rely on an obscure URL as the only control. Review what the dashboard exposes, disable unnecessary administrative functions, and stop the tunnel when remote access is no longer required.
Consider separating the gateway’s request endpoint from its analytics and administration interface. They can have different users, authentication requirements, availability expectations, and exposure policies. If the gateway combines them, verify that administrative routes cannot be reached by ordinary API clients and that dashboard users cannot retrieve raw sensitive content unless specifically authorized.
Build dashboards that lead to decisions

A useful dashboard should move from broad change detection to a focused workload and then to evidence for a routing decision. Start with trends, allow filtering by application and task, and preserve a path to conversation or request details for authorized investigators.
Recommended views
- Usage overview: Requests, logical tasks, conversations, input usage, output usage, latency, errors, and cost where reliable pricing data exists.
- Model distribution: Requested and selected model usage by application, task category, team, user type, and agent.
- Task analysis: Task volume, model mix, completion outcomes, turn counts, and usage per completed task.
- Conversation analysis: Distribution of turns, retries, provider attempts, and total time from first request to completion.
- Candidate overuse view: Routine task cohorts using capable models, accompanied by quality coverage and approved-alternative status.
- Actor concentration: Users, services, or agents responsible for unusual changes, without turning the view into a performance ranking.
- Routing audit: Requested-versus-selected model, fallback use, route-rule version, and changes over time.
- Telemetry health: Missing identifiers, unknown tasks, absent usage data, duplicate events, delayed outcomes, and schema-version distribution.
Every chart should identify its unit. “Requests per user” can mean application actions, gateway calls, provider attempts, or agent subcalls. Labeling the scope prevents incorrect comparisons. Filters should also preserve sample size. A dramatic percentage from a handful of events should not be presented with the same confidence as a stable, well-instrumented cohort.
Verify locally before exposing the dashboard
Start the gateway, telemetry components, and dashboard according to their own official documentation. Because no specific open-source gateway or dashboard package is defined in this workflow, we cannot provide truthful universal installation commands, default ports, credentials, file paths, or environment variables. Those values vary by implementation and must be obtained from the selected project.
Before adding remote access, verify from the machine running the dashboard that its local URL loads, authentication works, expected charts contain controlled test events, filters return the correct task and model groups, and unauthorized users cannot open restricted views. Also confirm whether the service listens only on loopback or on another local interface, then use that actual address and port as the tunnel target.
Perform an end-to-end telemetry test with a synthetic task. Confirm that the application reference, pseudonymous actor, task category, conversation, turn, requested model, selected model, timing, usage, and outcome appear as expected. Then test one retry and one multi-turn workflow to make sure totals roll up correctly.
Make the dashboard remotely reachable with Localtonet

After the dashboard works locally, a Localtonet HTTP tunnel can provide a public HTTPS address for it. The Localtonet client on the device establishes an outbound connection to our relay server, so this workflow does not require inbound router port forwarding, firewall changes, VPN setup, or a public IP address.
The tunnel is the connectivity layer only. Your gateway or analytics application continues to generate task classifications, token measurements, user associations, model-routing records, and overuse insights. We do not infer those semantic AI analytics from tunnel traffic.
Install and run the Localtonet client
Install the Localtonet application for the operating system on the device that can reach the dashboard. Keep the client running while remote access is required. Installation details can change by operating system and client version, so use the current Localtonet download and documentation flow rather than an unverified command.
Authenticate or select the device
Use the device-specific authentication token through the supported Localtonet workflow, then select the device that will run the tunnel. Never publish, embed, or reuse a guessed token. Treat the token as a credential.
Select an available relay server
Choose a currently available server or region from the Localtonet dashboard. Available values can vary, so obtain them from the current product rather than copying a hardcoded server code from an article.
Create an HTTP tunnel to the local dashboard
Configure the HTTP tunnel with the local IP address and port that you verified for the dashboard. HTTP tunnels can use Random Sub Domain, Custom Sub Domain, or Custom Domain process types, and each serves the target at a public HTTPS address. Custom-domain DNS requirements should be checked against the current documentation before configuration.
Start the tunnel and test the assigned address
Creating a tunnel does not start it. Use the Start button, then open the assigned public HTTPS address from a separate network or approved remote device. Confirm that authentication appears before sensitive dashboard content and that authorized views work correctly.
Stop or delete access when it is no longer needed
Stop the tunnel to make the public endpoint unavailable, or delete it when the configuration is no longer required. The tunnel is available only while the selected client device is connected and the tunnel is running.
For the current dashboard workflow and field names, consult the Localtonet HTTP tunnel documentation. Exact options can vary by plan, client version, region, or deployment, so verify current availability in the dashboard.
Choose the correct local target
The tunnel target must be the address and port where the dashboard is actually reachable from the Localtonet client device. Do not assume a default port. If the dashboard runs in a container, virtual machine, or another host on the local network, verify that the client device can connect to that address directly before troubleshooting the tunnel.
If only the dashboard should be exposed, target the dashboard service rather than a broader administrative interface. Where the application supports separate listeners, keep management, database, telemetry ingestion, and provider credentials off the public route. Expose only the minimum web surface required for the approved remote workflow.
Troubleshoot observability and remote access
The dashboard loads, but model or task data is missing
This is normally an instrumentation or analytics problem rather than a tunnel problem. Verify that the application supplies correlation context, the gateway emits routing and usage fields, the telemetry pipeline accepts the current schema version, and the dashboard’s time range includes the test event. Check unknown-category and rejected-event counters before assuming there was no traffic.
Request totals are much higher than expected
Determine whether the chart counts logical application requests, gateway calls, downstream provider attempts, agent subcalls, or retries. Inspect one known conversation and trace its identifiers across all events. A single user action may legitimately create multiple model calls, but those calls should roll up to the same task or agent run.
Cost totals disagree with provider billing
Check for missing usage records, duplicate events, retries, time-zone boundaries, cached responses, uncounted model calls, and outdated pricing data. Provider billing may also use commercial terms that differ from public prices. Label calculated cost as an estimate unless it is reconciled through the organization’s authoritative billing process.
A model appears overused, but quality data is unavailable
Treat the cohort as a candidate only. Add a suitable outcome signal and run a controlled comparison before changing routing. If quality cannot be measured directly, document the proxy being used and its limitations. Lower token usage does not establish equivalent output.
The dashboard works locally but not through the tunnel
Confirm that the selected Localtonet device is connected, the tunnel has been started, and the configured local target matches the verified dashboard address and port. Test the target from the same device running the Localtonet client. If that device cannot reach the service locally, the tunnel cannot forward requests to it.
The public address opens an error page
Check whether the local service expects a particular host configuration, path, or application base URL. Those behaviors belong to the dashboard application and should be configured using its official documentation. Do not change unrelated tunnel settings without first confirming that the application accepts requests arriving at the configured local listener.
The tunnel worked and then became unavailable
A Localtonet tunnel is available only while the selected client device is connected and the tunnel is running. Verify the device, client process, local dashboard, network connectivity, and tunnel status. If the device sleeps, restarts, or loses connectivity, remote access can stop until the required components are available again.
Remote users can see more information than intended
Stop the tunnel while reviewing the issue. Then inspect application authentication, role permissions, dashboard filters, raw-event access, exports, and administrative routes. The corrective control belongs primarily in the application’s authorization model. Re-enable access only after testing with a least-privileged account.
Frequently asked questions
Does Localtonet detect AI model overuse?
No. The AI gateway, application, and telemetry pipeline must generate model, user, task, conversation, turn, cost, and outcome data. With Localtonet, we can provide remote connectivity to the locally running gateway or dashboard. Tunnel connectivity and semantic AI observability are separate responsibilities.
Can token counts alone identify an unnecessarily capable model?
No. Token counts show consumption, not task difficulty or required quality. Evaluate task category, model choice, latency, conversation turns, retries, completion outcomes, and application-specific requirements. Use token data to locate candidates, then validate an alternative with representative tests.
Must the observability system store prompts and responses?
Not necessarily. Many model-usage questions can be answered with pseudonymous actor references, task categories, model names, usage measurements, turn counts, routing metadata, and outcome signals. Content should be collected only when necessary and should receive suitable redaction, access control, retention, and organizational review.
How should agent calls be counted?
Preserve every model attempt for operational debugging, but associate all related calls with a common agent run, conversation, and task. This supports both provider-call analysis and task-level cost analysis without pretending that every internal call represents a separate user action.
Does an HTTP tunnel make the dashboard secure automatically?
An HTTP tunnel provides a public HTTPS address and avoids inbound router port forwarding, but the dashboard still needs appropriate application authentication, authorization, least-privilege access, and safe data-handling controls. A public URL should never be treated as sufficient authorization by itself.
Why is the Localtonet dashboard address unavailable after it worked earlier?
The tunnel is available only while the selected Localtonet client device is connected and the tunnel is running. Also verify that the local dashboard remains active at the configured address and port. Device sleep, a stopped client, a stopped tunnel, service restarts, or network loss can interrupt access.
Should every simple task be routed to the least expensive model?
No. Model selection should account for quality, reliability, structured-output behavior, context requirements, latency, safety, and total task completion. The goal is an appropriate model for the workload, not the lowest per-request price regardless of outcome.
Access your self-hosted AI observability dashboard with Localtonet
Instrument and verify the AI analytics locally first, then use a Localtonet HTTP tunnel to make the approved dashboard reachable through a public HTTPS address without inbound router port forwarding.
Get Started Free →