Invoices, claims and identity documents are the most sensitive paper a company handles. KognitCapture runs on your servers, reads pages with engines inside the application, and makes every outbound connection something you configured.
This page is written for the person who has to sign off. It states what the software does, what it leaves to you as the operator, and what it does not do — without a badge standing in for an answer.
One JAR, one PostgreSQL database, one filesystem. This page covers what you install, how it scales, what it connects to, and which controls are yours to operate.
The deciding question for regulated documents is where they are processed. Here the answer is: on infrastructure you control, under your own security regime, with no vendor in the data path.
Role-based access with per-project hub grants, tenant and organisation isolation in code and on disk, encrypted secrets, account lockout, a central activity log, field-level history and retention policies.
02 · Yours
What you control as operator
Where it runs, who reaches it, TLS at your reverse proxy, disk and database encryption, backups, and whether any page is ever sent to a remote AI provider.
03 · Boundaries
Where the product stops
No vendor certification, no multi-factor sign-in or single sign-on yet, no database row-level security, no automatic redaction before model calls. Details below.
Data residency
Residency is a deployment decision, not a contract clause.
There is no KognitCapture cloud that your documents pass through. Where the data lives is decided by where you install the software and which endpoints you configure.
Recognition stays in the process
Tesseract OCR, the machine-learning classifiers and GGUF text models run inside the application. Vision models run in a local model server managed by the platform on the same host.
no external OCR service · no external inference service
Air-gapped operation
A documented installation recipe covers a network with no internet access: language packs and models are uploaded by hand instead of downloaded.
offline install · manual pack and model upload
Remote AI is opt-in, per organisation
Nothing is sent to a remote model provider unless an administrator configures that provider and a workflow step uses it. Local models can be switched on or off at platform and organisation level.
provider per organisation · step-level choice
No calls home
There is no licence server and no entitlement check to reach. Interface assets are served by the application itself — no content delivery network, no third-party fonts or scripts.
no licence server · all assets local
Deployment
One JAR. One database. Your infrastructure.
The application is a single self-contained Java archive that runs as a system service behind your reverse proxy.
We do not ship container images or Kubernetes manifests today, and there is no hosted edition. If you want it hosted, Automatize quotes that as a separate service.
PlatformLinux x86-64 — the supported, self-contained platform, with OCR libraries bundled
RuntimeJava 25, Spring Boot 4, single JAR, system service
DatabasePostgreSQL; schema created and upgraded by versioned migrations on start-up
FilesOn a filesystem, in a directory tree per tenant, organisation and project — never in the database
TLSTerminated at your reverse proxy
Memory4 GB heap for a small installation, 8 GB for real throughput; local models need additional memory outside the heap
ScalingSingle node, or several nodes over one database and one shared filesystem — no message broker
Node duties and capabilities
In a cluster, each node is told what it does. Keep the interface responsive on one node while two others do nothing but OCR.
Drain a node before maintenance. Work that was interrupted is picked up again.
drain · resume · interrupted-work recovery
Self-verification
Built-in end-to-end scenarios exercise the running installation, so an upgrade is checked by the product itself and not only by hope.
scenarios · dashboard · optional nightly sweep
Sealed configuration exports
Move a project or workflow between environments as an encrypted file bound to an installation key.
AES-256-GCM · HKDF-SHA-512
Access control
Deny by default. Grant by project.
Every API request must be authenticated. What a person can do follows from a role; which work they can open follows from hub grants.
Per-project roles beyond hub grants are enforced on the AI-agent (MCP) path today, not yet across the rest of the product.
Role
Scope
Can do
Superadmin
Platform
Tenants, system settings, cluster, every organisation
Tenant admin
Own tenant
Everything an admin can, in any organisation of the tenant
Admin
Own organisation
All designer rights, plus organisation record, branding, activity log and plugins; holds all hubs
Designer
Own organisation
Projects, templates, workflows, scripts, classifiers, datasets, connections, AI providers, reports, users, API keys. No hub unless granted
Verifier / User
Granted hubs
The work their hub grants allow, nothing in the designers
Hub grants
Scan Hub, mobile capture, Processing Hub and Verification Hub are granted separately, for one project or for all.
scan · mobile · processing · verification
Sign-in protection
Passwords are hashed with bcrypt. Five failed attempts lock an account for fifteen minutes; sign-in is rate-limited per address; an administrator can force a password change.
bcrypt · lockout 5 / 15 min · 10 attempts per minute per IP
Sessions and tokens
Bearer tokens with an HTTP-only cookie for the browser. Refresh tokens rotate, reuse of an old token is detected, and sessions can be revoked.
Shown once, stored as a SHA-256 digest, linked to named projects and limited by scope and allowed origin.
digest only · scopes · origin allow-list
Isolation
Clients separated in the code and on the disk.
Tenant, organisation and project boundaries are enforced in the service layer for every object access, and mirrored in the storage layout.
Three-level scoping
The scope of a request comes from the signed-in identity, never from a value in the request body.
tenant → organisation → project
Separate directory trees
Files for different tenants never share a directory. All paths are built by one central path service.
nested storage per level
Script sandbox
JavaScript and Groovy scripts run sandboxed with time and memory limits. Java scripts are unsandboxed by design and reserved for trusted administrators.
sandboxed JS and Groovy · trusted Java
What isolation is not
Isolation is implemented in application code. It is not additionally enforced by row-level security in the database. If you need database-enforced separation between two clients, run two installations.
application-level · no database RLS
Encryption
Secrets encrypted by the application. Volumes encrypted by you.
We are precise about this, because "encrypted at rest" is often said loosely.
Encrypted by KognitCapture
Project credentials, AI provider keys and other stored secrets, with AES-256-GCM. Configuration exports are sealed the same way.
AES-256-GCM · credentials vault
Encrypted by your platform
Document files and the database are not encrypted by the application. Use volume or disk encryption and your database's own facilities.
disk / volume encryption · database encryption
In transit
TLS is terminated at your reverse proxy, with the certificates and cipher policy you already manage.
your proxy · your certificates
Outgoing PDFs
Exported PDFs can be password-protected with AES-256 and digitally signed with the organisation's certificate. Signatures carry no trusted timestamp or long-term validation data.
AES-256 password · PKCS#12 signature
Audit trail
Who did what, to which document, and when.
Several logs answer different questions. Each has its own retention policy, set per level.
Activity log
Every changing API call, plus sign-in, failed sign-in, access denied and password change.
creates · updates · deletes · logins · denials
Field amendment history
For each field: what the engine read, what a rule changed, what a person corrected, and who verified it.
raw · amended · verified by · when
Document and step logs
The path a document took through its workflow, with status history and a log per step.
status history · step log · processing log
API call log
Inbound and outbound calls, including calls to AI providers with cost and latency.
inbound · outbound · AI calls
Agent channel
Actions by AI agents through the MCP server are recorded separately — reads as well as writes.
agent reads · agent writes
Retention
Keep it as long as you must. Not longer.
Retention is set per project and executed by a nightly sweep. It ships switched off: nothing is deleted until you define a policy.
Three policy actions
Mark documents as deleted, remove the files while keeping the record, or remove everything.
soft delete · delete files · delete all
Log retention
Separate retention periods for each log, at platform, tenant and organisation level.
per log · per level
No legal hold
There is no legal-hold feature that exempts selected documents from a retention policy. Handle holds by excluding the project from the policy.
not available
Privacy regulation
The product documentation describes which controls map to GDPR, HIPAA and CCPA obligations, as an aid to your own assessment. It is not legal advice, and the software is not "certified compliant" with any of them — compliance is a property of how you operate it.
control mapping · your assessment
Hardening
Assume the document is hostile.
A capture platform opens files from strangers all day. Several guards exist for exactly that.
Outbound request guard
Addresses that connectors call are checked against server-side request forgery. Blocking of private network ranges is available and ships off, because many installations export to internal systems.
SSRF checks · optional private-range block
Decode limits
Limits on image and PDF decoding, and a compression-ratio check on archives, stop oversized and bomb files.
image limits · PDF limits · archive ratio check
Host-key pinning
SFTP connections pin the server's SSH host key.
SSH host-key pinning
Inbound webhooks
Verified with an HMAC signature and an IP allow-list.
HMAC · IP allow-list
Reviewed, with open findings
The code base is security-reviewed. The latest review lists no critical and no high findings; medium findings are tracked and worked through. Ask us where that list stands today.
0 critical · 0 high · mediums tracked
Shared responsibility
What you are responsible for.
Self-hosting gives you control and hands you duties. These are yours.
01TLS, network segmentation and who can reach the application at all.
02Encryption of the disks and of the database.
03Backups of the database and the file tree, and testing that they restore. The product has no built-in backup.
04Operating-system and Java patching, and applying KognitCapture updates.
05Deciding whether a remote AI provider may see your pages, and under which agreement.
06Defining retention policies. None is active until you create one.
What we do not claim
If one of these is a hard requirement, you should know now
01
No vendor certification
Automatize holds no ISO 27001, SOC 2, HIPAA or PCI-DSS attestation for KognitCapture, and we will not imply one with a badge. The software runs under your certified environment, not ours.
02
No MFA or SSO yet
Users sign in with a username and password. Multi-factor sign-in and single sign-on through OIDC, SAML or LDAP are on the roadmap, not shipped. Until then, put the application behind your own access gateway if you need them.
03
No password policy engine
The minimum length is eight characters. There are no complexity or expiry rules and no idle time-out setting.
04
No redaction before model calls
Personal data is not masked before a page is sent to a model. Use a local model for sensitive material.
05
No validated PDF/A
PDF/A files are written to the standard's rules and are not run through a third-party conformance validator. PDF/UA is not claimed.
06
No e-invoice sending
KognitCapture reads Peppol, Factur-X, ZUGFeRD and XRechnung documents. It is not an access point and does not validate against network rules.
Common questions
Security & deployment
01Does any document leave our network?
Not unless you configure it to. OCR, classification and local models run inside the installation. A document leaves only through an export goal you defined, or to a remote AI provider you configured for a specific step.
02Can it run with no internet connection at all?
Yes. There is a documented air-gapped installation. Language packs and model files are uploaded manually, and there is no licence server to contact.
03We require SSO. Is that a blocker?
It depends on how strict the requirement is. Single sign-on is not built in today. You can place the application behind an identity-aware proxy as an outer gate. If users must authenticate to the application itself through your identity provider, wait for the roadmap item or talk to us about timing.
04How are two clients on one installation kept apart?
By tenant and organisation scoping in the service layer on every object access, and by separate directory trees on disk. It is application-level isolation, not database row-level security. For clients that demand hard separation, a second installation costs you a server, not a second licence negotiation — ask us about multi-installation terms.
05What does an AI agent get access to?
Nothing until you switch the MCP server on, at platform, tenant and organisation level. Then an agent gets the tools its OAuth consent or API key scope allows, acts as an identity with roles, and leaves a separate audit record of every read and write. Sessions can be revoked at any time.
06Is there a hosted version?
No. KognitCapture is self-hosted software. If you want someone to run it for you, Automatize can host and operate an installation as a separately quoted service — in which case that installation is still dedicated to you.
07Do you support Docker or Kubernetes?
The supported deployment is a JAR running as a system service on Linux x86-64. We do not publish container images or manifests today. Nothing prevents you from packaging it, but that packaging is yours to maintain.