# Pappur — Private AI Document Indexer

> Pappur is a privacy-first document intelligence API for AI agents. It reads documents with local OCR, turns them into multilingual semantic vectors with source pages and positions, and helps agents find the right passage for a question. Documents never leave the organisation, and vectors stay in the customer's own database.

Pappur opens soon. Contact us through the website for early access or an enterprise install.

## What it does

- Reads PDF, PNG, JPEG, DOCX, TXT and JSON; scanned pages go through GLM-OCR, tables are kept.
- Splits text into context-aware pieces and embeds them (Qwen3-VL-Embedding-2B, 2048 dimensions).
- Turns questions into query vectors and optionally ranks the candidates from your database.
- Runs its models locally; no external AI service is called.

## Enterprise (on your own server)

- Contracts and official documents: scanned contracts, specifications and official letters are read with OCR; tables and stamp text are kept, and every answer cites its page and position.
- Fit for the public sector: needs no internet while running and calls no outside service; model and software versions are pinned and verifiable by SHA256.
- Privacy-law (KVKK) ready: personal data is not passed to third parties or abroad; documents are encrypted while processed and deleted afterwards.
- Confidentiality: text and vectors are stored only in your own database; access is limited by user and API key.
- On your own server or your own cloud account; no GPU needed.
- Semantic search and reviewer assistant: searches meaning, not words; in document review it returns the passage for each checklist item with its source. Multilingual, including Turkish.
- Lower AI costs: documents are indexed locally in advance, so the language model gets only the relevant passages; fewer tokens, faster and more accurate answers.
- Admin-set limits and queue: file, page and character limits, queue capacity and storage are set in the admin panel; usage is reported per user and API key.

## How it works

1. Send a document to the API.
2. Pappur extracts, splits and embeds the text; each piece comes with its page and position.
3. Store the vectors in your own database; Pappur keeps none.
4. Turn the question into a vector, search your database, optionally let Pappur rank the candidates.
5. Give the matched text and its source to your language model.

## Plans

- Enterprise: unlimited use on your own server; installation, training and support included. Contact us for a quote.
- Cloud: pay as you go. 1 token = 1 PDF or image page, or 10,000 characters of extracted text, or 10,000 query characters.

## Links

- [Agent guide: API, limits, pricing and errors](https://pappur.beyondend.dev/agent-guide.md)
- [OpenAPI specification](https://pappur.beyondend.dev/openapi.json)
- [Website](https://pappur.beyondend.dev/)

## Türkçe özet

Pappur, yapay zekâ ajanları için gizlilik odaklı bir belge zekâsı API'sidir: sözleşme ve resmi belgeleri yerel OCR ile okur, anlamsal aramaya hazırlar, her yanıtı sayfa ve konumuyla kaynak gösterir. Enterprise sürümü kurumun kendi sunucusunda, internete ihtiyaç duymadan ve KVKK uyumlu çalışır; belgeler önceden yerelde indekslendiği için dil modeline yalnız ilgili parçalar gider ve yapay zekâ maliyeti düşer.
