Skavio Processing API
Your first processing job
API base: https://www.skavio.eu/v1. Create a Skavio account and company profile in the company workspace, then create a read/write API key. The secret is shown once. A company receives 1,000 trial credits.
Upload the source, estimate its price, submit the job with a stable idempotency key, poll until terminal status, then download the results. Do not put API keys in frontend code.
export SKAVIO_API_KEY='YOUR_KEY'
curl -sS https://www.skavio.eu/v1/uploads \
-H "Authorization: Bearer $SKAVIO_API_KEY" \
-F 'files=@invoice.pdf'
# Copy the returned upload ID; create request.json:
{
"operation": "flow-extract",
"upload_ids": ["UPLOAD_ID"],
"fields": ["supplier", "invoice_number", "currency", "total"],
"instructions": "Normalize unambiguous dates to YYYY-MM-DD.",
"max_credits": 30
}
curl -sS https://www.skavio.eu/v1/estimate \
-H "Authorization: Bearer $SKAVIO_API_KEY" \
-H 'Content-Type: application/json' --data-binary @request.json
curl -sS https://www.skavio.eu/v1/jobs \
-H "Authorization: Bearer $SKAVIO_API_KEY" \
-H 'Content-Type: application/json' \
-H 'Idempotency-Key: invoice-run-2026-001' \
--data-binary @request.json
curl -sS https://www.skavio.eu/v1/jobs/JOB_ID \
-H "Authorization: Bearer $SKAVIO_API_KEY"
curl -fS https://www.skavio.eu/v1/jobs/JOB_ID/results/0 \
-H "Authorization: Bearer $SKAVIO_API_KEY" -o data.xlsxEndpoint reference
Download complete OpenAPI JSON
Import this contract into Postman, Insomnia or your OpenAPI client generator. Request bodies, field constraints and multipart upload schemas are generated from the implementation. Routes use the authenticated company scope; IDs from another company return 404.
Operations and supported files
flow-extract accepts PDF, images and UTF-8 text. Every input document produces one document row; line items use a separate worksheet. Output: XLSX, CSV and JSON. Fields use unique names starting with a letter and containing letters, numbers or underscores, maximum 64 characters and 30 fields. Metadata names items, warnings, evidence, review_required and human_review are reserved. Missing values are null. Evidence is checked against the source text; this is an aid to review, not a guarantee of correctness.
transcribe and meeting accept one audio or video source. Normal file operations accept one source; pdf-merge accepts 2–20 PDFs. Single jobs keep their existing source rules. For server-managed file batches, use POST /v1/batches; each batch item becomes a normal child job with independent billing/result status. Saved workflows currently support document extraction.
Batch processing
API v1.2 adds first-class batches without changing the existing /v1/jobs contract. A batch is a parent object containing ordinary child jobs, so one bad file does not cancel successful files or require you to build your own queue.
Use POST /v1/batches/estimate to quote the whole batch, then POST /v1/batches with one stable Idempotency-Key. Creation is atomic: either the full batch is reserved/enqueued or no child job is created and no credits are reserved.
One operation over many files
curl -sS https://www.skavio.eu/v1/batches/estimate -H "Authorization: Bearer $SKAVIO_API_KEY" -H "Content-Type: application/json" -d '{
"operation": "webp",
"upload_ids": ["UPLOAD_1", "UPLOAD_2", "UPLOAD_3"],
"options": {},
"max_credits": 90
}'
curl -sS https://www.skavio.eu/v1/batches -H "Authorization: Bearer $SKAVIO_API_KEY" -H "Idempotency-Key: images-webp-2026-10-02-001" -H "Content-Type: application/json" -d '{
"operation": "webp",
"upload_ids": ["UPLOAD_1", "UPLOAD_2", "UPLOAD_3"],
"webhook_id": "WEBHOOK_ID",
"webhook_mode": "batch",
"max_credits": 90
}'
With the shorthand upload_ids, each upload becomes one child job. This works for File Toolbox single-file operations, OCR, Transcribe, Meeting and Flow extraction. For Flow this means independent per-document results; use explicit groups when one child job should contain multiple documents.
Grouped sources
Use items when an operation needs multiple sources per child job. For example, two independent PDF merges:
{
"operation": "pdf-merge",
"items": [
{"upload_ids": ["PDF_A", "PDF_B"]},
{"upload_ids": ["PDF_C", "PDF_D", "PDF_E"]}
]
}
The same grouping model can be used with Flow when several uploaded documents should produce one combined extraction export. A request must use either upload_ids or items, never both.
Lifecycle and partial failures
Batch status is queued → processing → completed | partial | failed. partial means at least one child completed and at least one failed. GET /v1/batches/{batch_id} returns aggregate progress plus every child job with its normal job result metadata. Failed child jobs release their own reserved credits; successful items remain charged and downloadable.
POST /v1/batches/{batch_id}/retry-failed creates a new batch containing only failed items. Supply a new stable Idempotency-Key. Source retention still applies, so retry before the original uploads expire.
Webhooks and downloads
webhook_mode can be batch (default), items, or both. Batch callbacks use the existing signing secret and signature rules with event types batch.completed, batch.partial or batch.failed. Item mode produces the normal signed job.completed/job.failed callbacks for each child.
Download child results through their authenticated job result paths, or after the batch is terminal use GET /v1/batches/{batch_id}/results.zip. The ZIP contains successful results grouped by item and a manifest.json. Combined ZIP generation is limited to 1 GB; larger batches should download item results individually.
Limits
A batch contains up to 100 items. Child jobs count toward the existing company queue limit of 100. The current fair-share scheduler still processes one active job per company, so large batches are naturally serialized for that company while the platform can process jobs from other companies concurrently. max_credits applies to the complete batch quote before anything is reserved.
Operation options
Put settings in the options object. File Toolbox validates source dimensions, duration, page ranges and passwords before processing. Passwords are forwarded privately to the PDF worker; do not log request bodies.
| Operation | Example options |
|---|---|
pdf-extract / pdf-delete | {"pages":"1-3,5"} |
pdf-rotate | {"pages":"1-2","angle":90} · 90, 180 or 270 degrees |
pdf-protect / pdf-unlock | {"password":"user-supplied-secret"} · 1–128 characters |
pdf-watermark | {"text":"DRAFT","pages":""} · empty pages selects all |
pdf-number | {"start":1,"pages":""} |
image-edit | {"width":1200,"height":1200,"angle":0,"flip":"none","format":"webp","crop":[0,0,800,800]} · resize contains image within the box and preserves aspect ratio; supply both dimensions |
image-compress | {"format":"webp","level":"balanced"} · formats jpg/png/webp; level balanced/strong |
audio-trim / video-trim | {"start":10,"end":60} · seconds within source duration |
video-rotate | {"angle":90} |
video-resize | {"height":720} · 240, 360, 480, 720, 1080, 1440, 2160 |
audio-compress / video-compress | {"level":"balanced"} |
transcribe | {"output_format":"txt"} · TXT, DOCX, SRT, VTT, JSON with native word timestamps and anonymous speaker segments |
Without a saved workflow, set operation, upload_ids, optional fields and instructions. With a saved workflow, also set workflow_id; the saved operation, options, fields and instructions take precedence. max_credits rejects a job if its quote is too high.
Job lifecycle, retries and idempotency
queued → processing → completed | failed. A job ID is returned immediately with HTTP 202. Poll every 3–10 seconds; back off on 429 or transient 5xx errors.
Use a unique Idempotency-Key for each business request, 8–128 characters: letters, numbers, underscore, dot, colon or dash. Retrying the identical request with the same key returns the same job without charging again. Reusing a key with different parameters returns 409. On a lost submission response, retry with the original key; do not create another key.
Failed jobs release reserved credits. Submit a new job with a new idempotency key when you intentionally retry a terminal failure. Temporary provider outages keep the job queued with its reservation. Durable checkpoints and retry schedules survive restarts; a replay with the same idempotency key does not create another charge. Permanent failures release reserved credits.
Authenticated downloads use the paths in result.files; no public result URL or API key in a query string. Use the dashboard to correct extracted scalar fields and regenerate exports without another processing charge. Review changes are recorded; line items must still be checked against the original. Results expire after 7 days. Preserve them in your system before expiration.
Credits, estimation and payments
POST /v1/estimate returns the price without starting work. The current contract reserves and charges that fixed quote on success. A failed job releases the full reservation. Estimates use server-measured source bytes, PDF pages and media duration.
File operations: 2 credits per started 25 MB or 10 pages (larger basis per source). OCR text: 3 per page; OCR DOCX/searchable PDF: 10 per page. Flow: 25 per page. Transcribe: 8 per started minute. Meeting: 15 per started minute. Media tools: 4 per started minute. Text extraction counts one page per started 4,000 source bytes. Image sources count as one page. Encrypted PDFs report 0 pages on upload; pdf-unlock validates your password and measures the real page count before quoting. Other operations require an unlocked source. Long document text is limited to prevent unexpectedly expensive extraction.
Monthly company plans: Starter €29 for 30,000 credits, Business €79 for 85,000 and Pro €199 for 225,000 and Enterprise €499 for 600,000. Credits are added after paid invoices and roll over. Manage or cancel through the billing portal. Each plan currently uses the same processing limits; priority service is not advertised.
Paid credit packs do not expire. Bonus packs change the effective euro price per credit. The dashboard shows credit units, not a refundable cash balance. All web tools and API now use the same prepaid wallet. Company admins can start a Stripe card checkout and reconcile a completed payment; signed Stripe webhooks reconcile it automatically. A checkout being opened does not add credits.
Keys, company scope and teams
Send Authorization: Bearer sk_live_…. Keys are hashed at rest and can have read-only or read/write scope. Administrators create and revoke keys through an interactive same-origin web session; API keys cannot create other keys or start billing checkouts. Revoke and replace keys to rotate them. Maximum 20 active keys per company.
API keys belong to a company and the member who created them. Removing membership invalidates access through that key. Team invitations are created in the dashboard. Invitations are accepted by the matching signed-in account; no email is sent automatically. Each account belongs to one company. Maximum 20 members. Company billing administrators manage settings; ordinary members can process and download files within their company. Do not share accounts between unrelated companies.
Limits, errors and retention
| Limit | Current release |
|---|---|
| Requests | 120 per minute per key/session |
| Upload batch | 100 files; total 500 MB |
| Company source storage | 2 GB; delete unused uploads to free it |
| Document | 100 pages per PDF; images count as one page |
| Media duration | Transcribe/Meeting: 10 hours; File Toolbox media: 2 hours |
| Queue | 100 pending jobs per company |
| Processing | 3 platform jobs across this deployment; 1 active per company for fair sharing |
| Results and source retention | 7 days; accounting metadata retained separately |
Errors return JSON with a detail field. 401: missing/invalid key or login. 402: insufficient credits. 403: scope/role/suspension. 404: object absent or not in your company. 409: incompatible state, mismatched idempotency or max-credit rejection. 410: expired result. 413: upload/storage limit. 422: invalid parameters/source. 429: request/queue limit. 502/503/504: processing/provider temporarily unavailable or timeout.
Delete a terminal job to remove platform result data while preserving accounting. Delete an unused upload to remove its source. In-flight sources cannot be deleted. Upstream processors and backups have their own retention; deletion is not a promise of immediate purge from all backups. Text from extraction is sent to the configured OpenAI model; Meeting and transcription use their existing configured processing providers. Review the privacy page and contact us for specific retention arrangements before uploading regulated data.
Webhook status
Register a public HTTPS callback URL in the dashboard or through an administrator session at POST /v1/webhooks. Save the returned signing secret; it is shown once. Pass its id as webhook_id when submitting a job.
On completion or failure we POST a JSON event with id, type (job.completed / job.failed), created_at and a job object. Job result paths still require your API key. No private source data is included in the event.
Verify X-Skavio-Signature as v1=HMAC_SHA256(secret, timestamp + "." + raw_body), using the exact body bytes and X-Skavio-Timestamp. Use constant-time comparison and reject timestamps older than five minutes. Deduplicate by X-Skavio-Event-Id. Acknowledge with a 2xx response after durable receipt.
Delivery is at least once. Up to six attempts are made, with backoff approximately 60, 120, 240, 480 and 960 seconds after failures. A disabled webhook receives no further deliveries. URL redirects are not followed; DNS is checked on every attempt and the HTTPS connection is pinned to a public IP. View status and errors under Webhooks.
import hashlib, hmac, time
def verify(secret, timestamp, raw_body, signature):
if abs(time.time() - int(timestamp)) > 300:
return False
expected = "v1=" + hmac.new(secret.encode(),
timestamp.encode() + b"." + raw_body, hashlib.sha256).hexdigest()
return hmac.compare_digest(expected, signature)Python example
import os, time, uuid, requests
base = "https://www.skavio.eu"
headers = {"Authorization": "Bearer " + os.environ["SKAVIO_API_KEY"]}
with open("invoice.pdf", "rb") as source:
r = requests.post(base + "/v1/uploads", headers=headers,
files={"files": ("invoice.pdf", source)}, timeout=180)
r.raise_for_status()
payload = {"operation": "flow-extract",
"upload_ids": [r.json()["uploads"][0]["id"]],
"fields": ["supplier", "invoice_number", "total"]}
q = requests.post(base + "/v1/estimate", headers=headers, json=payload, timeout=30)
q.raise_for_status()
payload["max_credits"] = q.json()["estimated_credits"]
# Persist this key with your business request before submitting.
job_headers = {**headers, "Idempotency-Key": str(uuid.uuid4())}
r = requests.post(base + "/v1/jobs", headers=job_headers, json=payload, timeout=30)
r.raise_for_status()
job_id = r.json()["id"]
for _ in range(720):
r = requests.get(base + "/v1/jobs/" + job_id, headers=headers, timeout=30)
r.raise_for_status()
job = r.json()
if job["status"] in ("completed", "failed"):
break
time.sleep(5)
else:
raise TimeoutError("Job still processing; resume polling with the same job ID")
if job["status"] == "failed":
raise RuntimeError(job["error"])
result = job["result"]["files"][0]
r = requests.get(base + result["download_path"], headers=headers, timeout=180)
r.raise_for_status()
with open(result["name"], "wb") as output:
output.write(r.content)JavaScript / Node.js example
// Node.js 22+; server-side only.
import { readFile, writeFile } from 'node:fs/promises';
import { randomUUID } from 'node:crypto';
const base = 'https://www.skavio.eu';
const auth = { Authorization: `Bearer ${process.env.SKAVIO_API_KEY}` };
async function json(path, options = {}) {
const r = await fetch(base + path, { ...options,
headers: { ...auth, ...options.headers } });
const data = await r.json();
if (!r.ok) throw new Error(JSON.stringify(data));
return data;
}
const form = new FormData();
form.append('files', new Blob([await readFile('invoice.pdf')]), 'invoice.pdf');
const uploaded = await json('/v1/uploads', { method: 'POST', body: form });
const payload = { operation: 'flow-extract',
upload_ids: [uploaded.uploads[0].id], fields: ['supplier','total'] };
const opts = { method: 'POST', headers: { 'Content-Type':'application/json' },
body: JSON.stringify(payload) };
const quote = await json('/v1/estimate', opts);
payload.max_credits = quote.estimated_credits;
const idem = randomUUID(); // Persist before submission.
let job = await json('/v1/jobs', { ...opts,
headers: { ...opts.headers, 'Idempotency-Key': idem },
body: JSON.stringify(payload) });
while (!['completed','failed'].includes(job.status)) {
await new Promise(resolve => setTimeout(resolve, 5000));
job = await json('/v1/jobs/' + job.id);
}
if (job.status === 'failed') throw new Error(job.error);
const file = job.result.files[0];
const r = await fetch(base + file.download_path, { headers: auth });
if (!r.ok) throw new Error('Download failed');
await writeFile(file.name, Buffer.from(await r.arrayBuffer()));Versioning and changelog
v1.2.0: first-class batch estimate/submit/progress, grouped sources, partial-failure accounting, failed-item retry, batch webhooks and aggregate ZIP downloads.
v1.0.0: company wallet and scoped keys, asynchronous file/audio jobs, document extraction with saved workflows and XLSX/CSV/JSON export, estimates, authenticated downloads and developer contract.
The API lives under /v1. Clients should ignore unknown response fields. Incompatible changes require a new major API version or an announced migration. The OpenAPI document and capabilities endpoint describe the currently deployed release.