How Pravia Works
What happens between a user's question and the answer, the numbers we measured, and what the system deliberately does not do.
What Pravia is
Pravia is a platform of AI employees that answer questions from the company's own documents. The employee works where it is needed: on the company website inside an embeddable widget, in Telegram, or through the API.
Video walkthrough
The same material as the text below, in video form.
Pravia walkthrough
Loading the walkthrough…
The walkthrough file in MP4 format.
Download MP4Three versions of one recording
How a question travels
Six stages from an uploaded document to a finished answer. First in plain language — how it works — and then in a clearly marked technical block — how exactly it is configured.
A block for engineers: real names, numbers and files. Everything else on this page describes the same system in plain language.
Text extraction
A document or a website is turned into text. PDF, DOCX, Markdown and TXT are accepted, as well as a website address and its sitemap. The server accepts these formats only — everything else is rejected.
Technical details
- allowed extensions: .pdf .md .markdown .docx .txt
- website import: page address and sitemap.xml
Splitting into pieces
Long text is split into medium-sized pieces so that search returns exact sections instead of whole documents. Neighbouring pieces overlap partially, so an answer does not break at the boundary.
Technical details
- RecursiveCharacterTextSplitter
- maxTokens 500, overlap 50
- lib/ai/chunk.ts
When text becomes numbers
Every piece becomes a compact numerical description of its meaning. It is computed by a model running on Pravia's own server: company text never leaves and is not sent to an external service.
Technical details
- onnx-community/embeddinggemma-300m-ONNX
- 768 dimensions, mean pooling + L2 normalisation
- ONNX Runtime in the server process
- lib/ai/embed.ts, lib/ai/models.ts
Storage
The text and its numerical description live in the same database as the rest of Pravia's data, and a dedicated index makes meaning search fast. The AI employee's identifier is duplicated directly onto the piece rows — deliberately, for retrieval speed.
Technical details
- PostgreSQL 16 + pgvector
- vector(768) column
- HNSW idx_chunks_embedding, vector_cosine_ops, m=16, ef_construction=64
- bot_id is denormalised onto piece rows on purpose, for retrieval speed
Retrieval
The user's question is matched by meaning against the pieces of this AI employee. At most three best matches are returned, and only those that pass the relevance threshold. If nothing matches by meaning, a keyword search is used instead.
Technical details
- <=> operator — cosine distance
- limit 3, threshold 0.7 — VECTOR_SEARCH constants in lib/ai/query-build.ts
- the candidate pool grows with corpus size, capped at 60
- per-document diversification: how many pieces one document may contribute
- rerank (bge-reranker-base): present but disabled by default — RERANK_ENABLED is unset
- keyword fallback: literal match (ILIKE)
Answer generation
The retrieved pieces and the question go to the model, which writes the answer. The answer arrives character by character and the client renders it as it comes in. The settings are chosen so the wording stays as predictable as possible.
Technical details
- DeepSeek-compatible chat completions
- temperature: 0, max_tokens: 8000 — a cap on answer length
- raw SSE streaming, no AI SDK
- <knowledge_context> marks the retrieved text as untrusted data and instructs the model to ignore any instructions inside it — prompt-injection isolation
- KNOWLEDGE_CONTEXT_WRAPPER in lib/ai/generate.ts
Retrieval quality — measured, not promised
Below are numbers from an automated run in the production configuration. We publish them together with the corpus and the command, because a figure without them means nothing.
- Measurement file
- tests/fixtures/rag-baseline.json · metricsVersion 1.1.0
- Measured on
- 29 September 2026
- Commit
- 272755f
- Configuration
- { limit: 3, threshold: 0.7 }, vector mode — this is Pravia's production configuration
- Corpus
- 67 documents / 717 pieces / 2 AI employees
- Scored cases
- 29 of 32: three have no ground truth, so they count neither as a hit nor as a miss
- recall@1
- 0.454
- recall@3
- 0.672
- precision@3
- 0.299
- MRR
- 0.644
What is not computed here
Two metrics are not computable in the production configuration, and we do not substitute other numbers for them.
- recall@5the production configuration returns at most 3 pieces, so "recall@5" would be recall@3 under another name
- nDCG@10same reason: measuring 10 positions against 3 results would just be nDCG@3 renamed
This measures retrieval only
The generation model was never called. We measured which pieces are found for a query, not which answers they produce.
About this corpus
This is the evaluation corpus, not the five-document demo set shown in the video.
What the system does not do
What follows are deliberate boundaries, not unfinished work. We list them so you can decide yourself whether that trade is acceptable.
No query rewriting
The question reaches retrieval exactly as it was typed. The system does not rewrite it, add synonyms, generate a hypothetical answer, or split a complex question into parts. The only transformation is a Russian-to-English translation, and it is switched off in the production configuration because the base model understands both languages anyway.
Technical details
- query rewriting, query expansion, HyDE, decomposition, spelling correction: none
- ru→en translation: conditional, switched off in the production configuration
No trimming of what was retrieved
The selected pieces are not cut down to fit a request-length budget. The figure 8000 in the settings is a cap on answer length, not a limit on how much text enters the request.
Technical details
- max_tokens 8000 — a cap on answer length
- trimming the retrieved context to a token budget: none
No citation verification
We ask the model to answer only from the retrieved pieces, but we do not check that every claim in the answer actually follows from them. That is a requirement in the instruction, not a technical guarantee.
Technical details
- verifying that the answer's claims follow from what was retrieved: none
- the basis is the system instruction text
Reranking is off by default
An extra pass by a model that re-scores the retrieved pieces exists in the system, but it does not run by default — one environment variable turns it on.
Technical details
- RERANK_ENABLED: unset by default → disabled
- reranking model: bge-reranker-base
The score in the UI is simplified
A number between 0 and 1 is shown next to a retrieved piece. For meaning search it is "how far apart", flipped into a convenient scale, not the probability of a correct answer. When the fallback result set is used, the number is always the same.
Technical details
- UI score = 1 − cosine distance
- fallback (keyword) results: score hardcoded to 0.5
A document is reloaded as a whole
Re-uploading a document never leaves half of it behind: the previous pieces are deleted and the new ones written in a single operation. If some pieces cannot be processed, that is recorded honestly as its own status rather than reported as a fully successful upload.
Technical details
- delete-then-insert inside one transaction
- unique index on (document_id, chunk_index) — the idempotency backstop
- partial run: embeddingStatus: "partial"
What it is built on
The actual stack, without marketing phrasing.
- Application
- Next.js 16 (App Router) + React 19
- Database
- PostgreSQL 16 + pgvector
- Background jobs
- BullMQ — optional: an empty REDIS_URL switches document processing to synchronous mode and the system keeps running
- Authentication
- Better Auth
- Styling
- Tailwind CSS 4
- Vectorisation
- @huggingface/transformers (ONNX), locally on the server
- Generation
- DeepSeek-compatible API, company-owned key (BYOK) on the Business plan
Ready to Hire Your AI Employee?
Start for free — no credit card required. Connect knowledge sources and deploy your AI employee in minutes.