# Pappur Agent API

> **Coming soon.** Pappur is not open for use yet: the API below is documented but returns `503 showcase_mode`. Contact us through the website for early access.

Pappur turns documents into searchable vectors for **your** database and turns questions into query vectors that find the right source pages. It does not store your documents or vectors and does not generate answers: your agent searches its own database and gives the retrieved text to its own LLM.

- Edition: **cloud**. Hosted in Oracle Cloud · Frankfurt. OpenAPI: `https://pappur.beyondend.dev/openapi.json`.
- Authentication: `Authorization: Bearer <PAPPUR_API_KEY>` (create keys in the Pappur console).
- Every `POST` needs an `Idempotency-Key` header (max 128 characters). Reuse it only to retry the identical request.

## 1. Embed a document

```bash
curl "https://pappur.beyondend.dev/v1/embed?filename=contract.pdf" \
  -H "Authorization: Bearer $PAPPUR_API_KEY" \
  -H "Idempotency-Key: contract-001" \
  -H "Content-Type: application/octet-stream" \
  --data-binary @contract.pdf
```

| Input | Processing | Measured as |
|---|---|---|
| PDF, PNG, JPEG | OCR (GLM-OCR) then text embedding; tables kept as HTML | real pages |
| DOCX, TXT | text read directly (no OCR), split into pieces of at most 2,000 model tokens on paragraph/sentence boundaries; headings are added to each piece as context | extracted characters |
| JSON `{"text": "..."}` | same as TXT; Markdown `#` headings become context | characters |
| JSON `{"items": [{"item_id","text","metadata","source"}]}` | your own pieces, embedded exactly as given (each at most 2,000 model tokens) | characters |

Encrypted, broken or over-limit files are rejected immediately, before queueing and without charge.

The response is `202` with a `job_id`. Each file is one job.

## 2. Follow the job

- `GET /v1/jobs/{job_id}?wait=60` waits up to 60 seconds and returns as soon as the status changes (long polling; no public address or webhook needed). Repeat until the status is final. `status`: `queued` → `running` → `success` | `failed` | `cancelled`; `error` holds the reason. Without `wait` the call returns immediately.
- `DELETE /v1/jobs/{job_id}` cancels your own job (refunded) or removes a finished result.
- Results expire 15 minutes after the job finishes.

## 3. Store the result

`GET /v1/jobs/{job_id}/result` returns `items[]`. Store each item in your vector database:

- `item_id`, `text` (the exact text that was embedded, including heading context), `embedding` (2048 normalized floats), `embedding_space`, `metadata` (`headings`, `table_html` for tables, OCR flags), `source` (`page` and normalized `bbox` [left, top, right, bottom] for OCR; `char_span` for text).
- Also returned: `page_count`, `warnings`, `usage` (pages or characters), `billing`, `exports`. `GET /v1/jobs/{job_id}/download/{name}` fetches a name from `exports`.
- Never mix vectors with a different `embedding_space` (current: `pappur:qwen3-vl-embedding-2b:q8_0:6a1b92741466:b11240:prompt-v1:2048`), even if the dimension matches.
- OCR can misread names and numbers; `metadata.ocr_verified` is `false` for OCR text.

## 4. Find sources for a question

```bash
curl "https://pappur.beyondend.dev/v1/query" \
  -H "Authorization: Bearer $PAPPUR_API_KEY" \
  -H "Idempotency-Key: question-001" \
  -H "Content-Type: application/json" \
  -d '{"query": "Who is responsible for site safety?"}'
```

1. `POST /v1/query` with `{"query": "...", "top_k": 5}` (at most 1,000 model tokens) returns `query_embedding` and `job_id` synchronously.
2. Search your own database with `query_embedding` (cosine / dot product).
3. Optional: `POST /v1/read/{job_id}/rank` with `{"candidates": [original Pappur items from your database], "top_k": 5}` returns `matches` sorted by `score`, with their text and source. One ranking per query; an identical retry returns the same result. You may also pass `candidates` directly in `/v1/query`.
4. Give the matched `text` and its `source` to your LLM.

`/v1/query` waits at most 150 seconds. When the query queue is busy, use the asynchronous `POST /v1/read` (same body, returns `202` + `job_id`) and follow the job with `?wait=60`.

## Limits

| Limit | Value |
|---|---|
| File size | 150 MB |
| PDF pages per job | 200 |
| Text characters per job (DOCX/TXT/JSON) | 1,000,000 |
| Query length | 1,000 model tokens |
| Document queue (waiting + running) | 30 jobs |
| Query queue | 300 queries |
| Document jobs at once | 1 per account (queries are not limited) |
| Upload time | 15 minutes |
| Job time limit | 5 h 10 min |

## Queue and retries

- Documents are processed one at a time, in order; queries run in parallel on a separate worker.
- Send one document at a time: wait (with `?wait=60`) until the previous job finishes (`success`, `failed` or `cancelled`) before sending the next file, otherwise you get `429 account_queue_full`.
- `GET /health` (no key) reports `queue.documents` and `queue.queries`: `waiting`, `running`, `capacity`, `level` (`idle`, `busy`, `crowded`, `full`), and for documents the queued `pages` and `characters`. Check it before sending large batches.
- `429` means nothing was accepted: `queue_full`, `account_queue_full` or `queue_storage_full`. Wait for `Retry-After` seconds and retry with the same `Idempotency-Key`.

## Pricing (Cloud)

1 token = 1 PDF or image page (OCR included) **or** 10,000 extracted characters of DOCX/TXT/JSON **or** 10,000 query characters.

- The amount is measured **once, when the upload is accepted**, and deducted then. Failed or cancelled jobs are refunded automatically.
- Query characters accumulate per account; each full 10,000 costs one token (a typical query is 100-200 characters). Candidate ranking is free.
- `max_charge_tokens` (optional on `/v1/embed`) caps one job; if the measured amount is higher, the upload is rejected with `402 budget_exceeded` and nothing is charged.
- `402 insufficient_balance` means the account cannot cover the measured amount. New accounts receive 10 trial tokens; packages are bought in an external sales system.

## Errors

| Code | Meaning | What to do |
|---|---|---|
| `unsupported_file` (415) | file type not supported | send PDF, PNG, JPEG, DOCX, TXT or JSON |
| `file_too_large`, `page_limit`, `character_limit`, `too_many_items` (413) | over a limit above | split the document |
| `encrypted_pdf`, `invalid_document`, `invalid_input`, `empty_file`, `no_text` (400) | unreadable input | fix the file |
| `chunk_limit` (400) | one of your JSON items exceeds 2,000 model tokens | split that item |
| `query_limit` (400) | query over 1,000 model tokens | shorten the query |
| `query_timeout` (504) | the synchronous query waited too long in the queue | use `/v1/read` and poll |
| `idempotency_key_required` (400), `idempotency_conflict` (409) | missing key, or key reused for different input | use a new key per input |
| `queue_full`, `account_queue_full`, `queue_storage_full` (429) | queue capacity reached | wait `Retry-After`, retry |
| `insufficient_balance`, `budget_exceeded` (402) | not enough tokens / above your cap | top up or raise `max_charge_tokens` |
| `job_not_found` (404), `result_not_ready` (409), `job_expired` / `result_expired` (410/404) | wrong job, still running, or expired | poll again or resubmit |
| `embedding_space_mismatch`, `invalid_corpus`, `invalid_vector`, `unnormalized_vector` (400) | ranking candidates are not current Pappur items | re-embed or send original items |
| `read_already_ranked` (409) | this query was ranked with other candidates | send a new query |
| `ocr_failed`, `processing_failed`, `timeout` | processing failed (refunded) | retry later |
| `invalid_api_key`, `authentication_required` (401) | bad or revoked key | create a new key |

## Privacy and models

Files are kept only while queued and processed, results for 15 minutes; then they are deleted. While held, every file is encrypted on disk with a per-job key that exists only in memory and is destroyed with the job, so nothing readable remains on disk. Pappur never stores your vectors and keeps only account, job and usage metadata. Inference runs locally on the Pappur server; no external AI service is called.

- OCR: GLM-OCR (F16 + Q8_0 vision), `ggml-org/GLM-OCR-GGUF@65a42de1148d`, sha256 `b06675e983db9593db78603b06f097e48c0cf078b37731c0a09612f4a249cf6f`
- Embedding: Qwen3-VL-Embedding-2B (Q8_0, llama.cpp b11240), `DevQuasar/Qwen.Qwen3-VL-Embedding-2B-GGUF@6a1b92741466`, sha256 `7552c2f699c546ce46abd6b66b2aa16ae667c88c830efbd352b12224d4613492`
