
Put a narrow, authenticated ingestion endpoint at the public boundary while keeping embeddings and vector storage private
Testing a local RAG pipeline with realistic events does not require publishing the vector database itself. A safer architecture accepts production-like HTTPS webhooks through a narrowly scoped local receiver, validates every event, records it durably, and moves expensive chunking and embedding work into an asynchronous queue. In this guide, we design that boundary, add signature validation, replay protection and idempotent indexing, verify retrieval quality, and then expose only the receiver through a Localtonet HTTP tunnel. The vector database remains bound to localhost or another private interface throughout the workflow.
π What's in this guide
Use an ingestion boundary, not a public vector database
A webhook provider needs an HTTP endpoint to which it can deliver events. It does not need direct access to a vector database, embedding service, worker control port, container dashboard, or retrieval API. Treating those internal components as webhook destinations unnecessarily expands the public attack surface and tightly couples an external event format to the storage layer.
The recommended boundary is a small HTTP receiver with one responsibility: decide whether an incoming event is authentic, timely, valid, authorized and safe to enqueue. Once accepted, the receiver returns promptly. Separate workers can then fetch approved source content, normalize it, split it into chunks, generate embeddings and update the private vector store.
Webhook producer
|
| HTTPS event
v
Public Localtonet URL
|
v
Local webhook receiver
| signature validation
| timestamp and replay checks
| schema and size validation
| durable event record
v
Private work queue
|
v
RAG ingestion worker
| content normalization
| chunking
| embedding generation
| idempotent replacement
v
Vector database on localhost or private network
This separation gives each layer a narrow trust boundary. The public receiver handles untrusted network input. The queue absorbs bursts and supports retries. The ingestion worker performs expensive or failure-prone processing. The vector database accepts traffic only from trusted local processes or private network peers.
The public destination should be the local HTTP receiver, not the database listener. Bind the database to localhost or an appropriate private interface, require its native authentication where available, and prevent direct public routing to its port. A tunnel is a connectivity mechanism, not a substitute for application authorization.
Define the threat model before accepting live events

A public webhook URL can receive requests from anyone who discovers or guesses it. HTTPS protects the network exchange between the sender and the public endpoint, but it does not prove that the request came from the expected webhook provider. Authentication and authorization must therefore be enforced by the receiver.
Consider at least five classes of failure. An attacker might forge an event, replay a previously valid request, send a very large payload, exploit a parser or document converter, or manipulate metadata so that indexed content crosses tenant or access-control boundaries. Ordinary delivery failures matter too: the provider may retry an event, events may arrive out of order, and an update may race with a deletion.
| Risk | Control at the receiver | Control behind the receiver |
|---|---|---|
| Forged webhook | Validate the provider's documented signature over the exact raw request body | Authorize the source, tenant and event type before indexing |
| Replay of a valid request | Require a signed timestamp and reject stale requests | Store the provider event ID with a uniqueness constraint |
| Duplicate delivery | Return a consistent success response for an already accepted event | Use deterministic document and chunk identities |
| Oversized or malformed input | Apply body-size limits, content-type checks and strict schema validation | Limit extracted text, chunk count and processing time |
| Out-of-order updates | Capture the source version or modification time | Reject an older mutation when a newer version is already indexed |
| Cross-tenant data exposure | Derive tenant identity from trusted configuration or validated claims | Preserve tenant filters in indexing and every retrieval query |
| Parser exploitation | Accept minimal JSON rather than arbitrary uploaded files where possible | Sandbox document conversion and keep parsers updated |
Sign the raw bytes, not reconstructed JSON
Many webhook schemes use a keyed message authentication code, often HMAC, but the exact algorithm, header name and signed message format are provider-specific. Follow the sender's documented scheme exactly. Capture the raw request body before parsing it because serializing the parsed object again can change whitespace, key ordering or character encoding and invalidate an otherwise correct signature.
Compare the expected and supplied signatures with a constant-time comparison function. Never log the shared secret or a complete authorization header. Store secrets outside source control and rotate them through a controlled procedure that can temporarily accept both the old and new secret if the provider supports overlapping rotation.
Combine freshness checks with durable deduplication
A timestamp window limits how long a captured request remains usable, but it does not stop multiple replays inside that window. A unique event identifier stops duplicates, but it does not by itself prove freshness. Use both controls when the sender supplies both signed values.
Choose the acceptance window according to expected clock skew and delivery behavior. There is no universal safe number. Monitor rejected timestamps before tightening the window, synchronize system clocks, and document whether delayed provider retries can legitimately fall outside it.
Design a minimal webhook event contract
Avoid placing an entire document in the event unless the source system requires that pattern and the data has been approved for the local test environment. A smaller event can identify the resource and mutation, after which an authorized worker retrieves the current content through a separate source API. This reduces webhook payload exposure and avoids indexing an obsolete body when several updates arrive close together.
A useful logical contract usually includes an event ID, event type, source identifier, resource identifier, source version and event time. It may include a tenant or workspace identifier, but the receiver must not blindly trust that field to grant access. Tenant mapping should come from authenticated source configuration or another trusted mapping.
{
"event_id": "provider-assigned-unique-id",
"event_type": "document.updated",
"source": "approved-source-name",
"resource_id": "source-document-id",
"resource_version": "source-version-or-revision",
"occurred_at": "provider-supplied-timestamp"
}
This is an illustrative internal shape, not a claim about any provider's payload. Adapt the mapping to the provider's documented schema and signature rules. Preserve the original event type and identifiers in the durable event record so that an investigation can trace an indexed result back to its trigger without retaining unnecessary sensitive content.
When practical, let the webhook say that a resource changed and let the worker fetch the latest authorized representation. This naturally coalesces several rapid updates and reduces the chance that an older event overwrites newer content. If the sender only supports full-content payloads, minimize retained fields and define a deletion schedule for raw event bodies.
Decide how each mutation changes the index
Create and update events normally converge on an upsert operation. A deletion event should remove all chunks associated with the logical document, not only a single vector. Permission-change events require special care because stale access metadata can expose content during retrieval even when the text and embedding remain correct.
Unknown event types should not silently enter the ingestion queue. Reject them or quarantine them according to a documented policy. A successful HTTP response should mean a specific thing, such as βauthentic event durably accepted,β rather than βthe entire embedding workflow completed.β
Build the local webhook receiver
The receiver should do little synchronous work. Signature verification, freshness validation, schema checks, deduplication and durable queue insertion belong on the request path. Content fetching, document conversion, chunking, embedding and vector writes generally do not.
Capture the raw request safely
Accept only the required route, method and content type. Enforce a body-size limit before parsing. Retain the exact raw bytes long enough to perform the provider's documented signature calculation.
Authenticate before processing content
Extract the documented signature and timestamp headers, calculate the expected signature with the configured secret, and compare values in constant time. Reject missing, malformed, stale or invalid authentication data.
Validate the event schema and authorization scope
Parse JSON only after authentication. Allow known event types and required fields, reject unexpected structures where practical, and verify that the configured source is permitted to mutate the referenced tenant or collection.
Claim the event ID atomically
Insert the event into a durable store with a uniqueness constraint on the source and event ID. If it already exists, return the response selected for an accepted duplicate without creating a second job.
Enqueue work in the same durable boundary
Use a transaction or transactional outbox pattern so the accepted event and pending work cannot diverge. Do not acknowledge success if the only copy of the job exists in process memory.
Return promptly and observe the result
Respond only after durable acceptance. Record a correlation identifier, status and processing latency without logging secrets or unnecessary document content. Let a worker process the queued job independently.
receive(request):
raw_body = read_with_size_limit(request)
authentication = extract_provider_headers(request)
if not verify_provider_signature(raw_body, authentication):
return unauthorized
if not timestamp_is_acceptable(authentication.timestamp):
return unauthorized
event = parse_and_validate_json(raw_body)
begin transaction
inserted = insert_event_if_absent(
source = configured_source,
event_id = event.event_id,
payload_reference = event.resource_id
)
if inserted:
insert_outbox_job(event.event_id)
commit transaction
return accepted
The pseudocode deliberately omits header names, algorithms and status codes because those are part of the sender's contract. Implement them from that contract rather than assuming that all webhook systems behave alike.
Keep status reporting separate from sensitive content
Useful logs include a generated request correlation ID, a hashed or truncated event identifier, event type, source, validation outcome, queue outcome and elapsed time. Avoid full payload logging by default. RAG content often includes proprietary text, personal data or credentials accidentally embedded in documents.
Health endpoints should report whether the receiver process is ready to accept work, but they should not reveal queue credentials, database addresses, collection names or internal exception traces. If a readiness check depends on storage, use a narrowly scoped check rather than exposing an administrative interface.
Make chunking and vector indexing idempotent

Receiver deduplication prevents the same event from creating multiple jobs, but it does not fully solve indexing consistency. A worker can crash after writing some vectors but before marking a job complete. Two distinct events can also refer to the same document version. The vector update itself must therefore be repeatable.
Use deterministic identities
Assign a stable logical document key from trusted source, tenant and resource identifiers. For each indexed revision, derive chunk identities from values that remain stable when the same version is retried, such as the document key, source version, chunking-strategy version and chunk position. A cryptographic digest can make the identifier compact, but the canonical input must be documented.
Do not derive identity only from chunk text. Identical paragraphs can legitimately appear in different documents, and permission metadata can differ even when text matches. Conversely, a minor formatting change can alter text hashes and leave obsolete chunks behind unless replacement is document-scoped.
Replace a document version as a unit
The safest available mechanism depends on the selected vector database. Where supported, write a new version and then atomically switch an active-version marker. Otherwise, upsert the new deterministic chunk set and remove chunks belonging to older versions only after the new set is complete. If neither operation can be atomic, retain an indexing-state record so retrieval can filter out incomplete versions.
Track the embedding model identifier, vector dimension, normalization choice, chunking strategy version and source revision in metadata. When any compatibility-critical value changes, build a separate index generation or collection rather than mixing incompatible vectors.
| Metadata | Why retain it | How it is used |
|---|---|---|
| Logical document key | Groups all chunks from one source resource | Replacement, deletion and traceability |
| Source version | Orders mutations and identifies stale work | Rejecting older updates |
| Chunk position | Preserves document order | Context reconstruction and deterministic IDs |
| Chunker version | Explains why boundaries changed | Controlled reindexing and comparisons |
| Embedding model version | Prevents incompatible vector mixtures | Index generation selection |
| Authorization attributes | Preserves retrieval restrictions | Mandatory query-time filtering |
A semantically correct result can still be unauthorized. If retrieval depends on tenant, workspace, document classification or user permissions, enforce those filters in every query path. Do not ask a language model to decide whether a retrieved chunk should have been visible.
Handle deletes, retries and poison jobs deliberately
Deletion should be idempotent: deleting an already absent document is normally a successful final state. A transient source API or embedding failure should schedule a bounded retry with backoff. A permanent schema, authorization or unsupported-format failure should move to a dead-letter or quarantine state with enough sanitized context for diagnosis.
Set a maximum attempt policy and expose queue age, failure count and oldest pending job as operational metrics. An HTTP receiver can appear healthy while indexing is hours behind, so request success alone is not a useful measure of pipeline health.
Verify the entire pipeline locally before creating a tunnel
Start with every component private. Run the receiver on a loopback or private address, keep the vector database on a separate local port, and verify that only the receiver route is intended for eventual publication. The exact startup commands, package names and database ports depend on the framework and vector store selected for the project, so they should come from those projects' current official documentation rather than being guessed.
Test authentication with exact request bytes
Build fixtures from the webhook provider's documented examples or from a development webhook environment. Test a valid signature, one altered body byte, a missing signature, an invalid encoding, an expired timestamp and a valid request replayed inside the freshness window. Confirm that only the first authentic event creates a job.
Test event ordering and worker recovery
Send two updates for the same resource in reverse order. The final active index must represent the newer source version. Terminate the worker after it writes some chunks, restart it, and verify that retrying converges on one complete chunk set without duplicates or stale active chunks.
Test retrieval quality, not only vector counts
An ingestion run is not successful merely because vectors were written. Maintain a small evaluation set of questions with expected source documents or passages. After each chunking or embedding change, run the same queries and compare retrieval results.
Also inspect operational behavior: acknowledgement latency, queue delay, worker duration, source-fetch errors, embedding errors and vector-write retries. These measures reveal bottlenecks that a small static test file may not expose.
Expose only the receiver with a Localtonet HTTP tunnel

Once local tests pass, Localtonet can provide the public HTTPS ingress path to the receiver. Our client runs on the machine that can reach the local service and establishes an outbound connection to a Localtonet relay server. This avoids inbound router port forwarding, a public IP requirement, VPN setup and inbound firewall changes.
The HTTP tunnel should target the receiver's local IP address and port. It should not target the vector database, queue administration interface, model server or container dashboard. Creating a tunnel does not make it active automatically. It must be started, and it remains available only while the selected client device is connected and the tunnel is running.
Install and run the Localtonet client
Install the current client for the operating system on the machine that can reach the webhook receiver. Use the current download and installation instructions shown by our platform because commands and package details can change by client version.
Authenticate or select the receiver device
Use the device-specific authentication token supplied through the Localtonet dashboard. Treat it as a secret, do not paste it into source code or logs, and never substitute a token from an example.
Select an available relay server
Choose a server or region currently offered in the dashboard. Available values can vary, so this guide does not hardcode a server code.
Create an HTTP tunnel to the local receiver
Enter the local IP address and port on which the receiver is listening. Select the appropriate HTTP process type from the options currently available, such as Random Sub Domain, Custom Sub Domain or Custom Domain. All three process types serve the content at a public HTTPS address, but availability may depend on current product configuration or plan.
Start the tunnel and register its public URL
Use the Start button, copy the assigned public HTTPS URL, and configure the webhook provider to call the exact ingestion route beneath that URL. Preserve the provider's signature settings and use a provider-supported test delivery before enabling broader traffic.
Stop or delete access when the test ends
Stop the tunnel when public ingress is no longer required. Delete it if the endpoint should not be reused, and remove or rotate the webhook secret according to the sender's lifecycle controls.
For the current interface and configuration sequence, consult our Localtonet HTTP tunnel documentation while setting up the tunnel. Exact client commands, server availability and custom-domain DNS requirements should be taken from the current dashboard and documentation rather than copied from an old environment.
The Localtonet client must be able to reach the receiver at the configured local IP address and port. If the receiver runs in a container or virtual machine, verify connectivity from the client environment itself. A service that responds on the host may not automatically be reachable through an isolated container network.
Keep signature verification, freshness checks, event authorization and input limits enabled after the tunnel starts. Do not rely on URL secrecy. Also verify that the receiver does not expose framework documentation, debug routes, stack traces or unrelated application endpoints.
Test the public boundary safely
Begin with the webhook provider's test-delivery function if one exists. Confirm that the request reaches the receiver, the signature validates against the untouched raw body, one durable job is created, the worker indexes the expected resource, and a retrieval query returns the intended passage under the correct authorization filter.
Then repeat a delivery and confirm that it does not create duplicate active chunks. Stop the Localtonet tunnel and verify that the public endpoint is no longer available. Restart it only when continued external testing is intended.
Operate the ingestion path without losing control of it
A realistic test can send production-derived data into a developer-controlled machine, which may create privacy, retention or policy obligations. Obtain approval before routing such events, minimize fields, use dedicated test tenants where possible, and define how long raw events, extracted text, vectors and logs will remain on the machine.
Monitor both ingress and eventual indexing
Track separate states for received, authenticated, accepted, queued, processing, completed, retrying, quarantined and deleted events. This makes it possible to distinguish a provider delivery problem from a worker or retrieval problem.
Useful alerts include sustained signature failures, increasing queue age, repeated processing attempts, a rising quarantine count and retrieval evaluation regressions. Avoid alert messages that copy document bodies or secrets into another system.
Use controlled replay for debugging
Replaying a stored request through the public endpoint may fail correctly because its timestamp is stale. For internal diagnostics, create a privileged local replay tool that re-enqueues a sanitized, previously authenticated event by its internal record ID. Keep that tool inaccessible from the public route and record who initiated the replay.
Plan schema and model migrations
Version the webhook mapping, normalized document schema, chunking strategy and embedding configuration independently. A webhook schema update should not force an immediate embedding migration, and an embedding migration should not change the sender's public contract.
For major index changes, build a new generation alongside the existing one, evaluate it, and switch retrieval only after it passes quality and authorization tests. Retain a rollback path until the new generation is proven.
Troubleshoot from the boundary inward
| Symptom | Likely area | Checks |
|---|---|---|
| Public URL does not respond | Tunnel lifecycle or local reachability | Confirm the client is connected, the tunnel is started, and the configured local target responds from the client device |
| Local request works but signature fails publicly | Signing implementation | Verify raw-byte handling, header parsing, encoding and the provider's exact signed-message format |
| Provider retries continuously | Receiver response or latency | Check response status, acknowledgement time, body-size limits and whether expensive work remains on the request path |
| Accepted events never become searchable | Queue or worker | Inspect outbox dispatch, queue age, worker errors, source authorization and vector-write state |
| Duplicate chunks appear | Index identity design | Review event uniqueness, deterministic chunk IDs, partial-write recovery and old-version cleanup |
| New content is indexed but retrieval is poor | Extraction, chunking or embeddings | Inspect normalized text, chunk boundaries, model compatibility, filters and evaluation-query results |
Frequently asked questions
Should I expose my local vector database to receive RAG webhooks?
No. Expose a narrowly scoped HTTP receiver and keep the vector database on localhost or a private network. The receiver authenticates and validates events, then a trusted worker updates the database. The webhook sender does not need database access.
Does HTTPS remove the need for webhook signature validation?
No. HTTPS protects the network exchange, but the receiver still needs to authenticate the sender. Implement the provider's documented signature scheme over the exact raw body, use constant-time comparison, and add replay controls when the provider supplies timestamps and event IDs.
Why should embedding work run asynchronously?
Fetching content, parsing files, splitting text, generating embeddings and writing vectors can be slow or temporarily fail. A durable queue lets the receiver acknowledge an authenticated event promptly, absorb bursts, retry transient errors and preserve a clear processing history.
How do I prevent duplicate vectors when a webhook is retried?
Deduplicate the provider event ID at acceptance time and make the vector update independently idempotent. Use stable document identities, deterministic chunk identities, source-version checks and document-scoped replacement or cleanup. This also protects against worker retries after partial failure.
What should the webhook payload contain?
Prefer a minimal event containing the event identity, mutation type, source resource identity, source version and timestamp. Let an authorized worker fetch the current document when practical. If full content must be delivered, minimize retained fields and apply explicit storage and deletion policies.
What happens when I stop the Localtonet tunnel?
The public tunnel is available only while the selected Localtonet client device is connected and the tunnel is running. Stopping it removes that active ingress path. The local receiver, queue and vector database can continue running privately if you leave those processes active.
Can Localtonet make an insecure webhook handler safe automatically?
No. Localtonet provides the public connectivity path to the configured local service. Your application must still validate signatures, reject replays, authorize event scope, limit input, protect secrets and handle failures safely. Publish only the receiver and stop the tunnel when external access is no longer needed.
Test your RAG ingestion boundary with Localtonet
Keep your vector database private, validate the receiver locally, and then use a Localtonet HTTP tunnel to give an approved webhook producer a public HTTPS destination. Start the tunnel only for the required test window and retain application-level authentication throughout.
Get Started Free β