Working with Files
A practical guide to the two file management models in the Haufe Files API: User Files and Space Files. Understand when to use each model and how to follow the end-to-end upload and consumption flow.
Working with Files
Haufe Files supports two distinct file models depending on who owns the content and how it is shared:
| User Files | Space Files | |
|---|---|---|
| Route prefix | /v1/files | /v1/spaces/{spaceId}/files |
| Ownership | Tied to a single user | Tied to a tenant Space |
| Sharing | Private to the user | Private (owner only) or Public (all tenant users) |
| Required headers | API key + x-user-id | API key + x-user-id + x-tenant-id |
| RAG scope | Per-file | Per-space (across all space files) |
How File Processing Works
Regardless of the model, every uploaded file goes through the same processing pipeline before it can be consumed:
- Upload — you request a pre-signed S3 URL from the API and PUT the file directly to S3.
- Antivirus scan — AWS GuardDuty scans the file. Infected files are quarantined and the file status is set to
FAILED. - Text extraction — the appropriate processor runs based on the file type (PDF, Word, Excel, PowerPoint, HTML, plain text / CSV / Markdown).
- Chunking & embedding — the extracted text is split into chunks, each one embedded via Amazon Bedrock and stored in the knowledge base.
- Ready — the processing status transitions to
PROCESSEDand the file is available for RAG or direct content retrieval.
Processing statuses
| Status | Meaning |
|---|---|
PROCESSING | Pipeline is running |
PROCESSED | File is ready to use |
FAILED | Processing or antivirus check failed |
Supported file types: PDF (.pdf), Word (.docx), Excel (.xlsx), PowerPoint (.pptx), HTML (.html, .htm), plain text (.txt), CSV (.csv), Markdown (.md)
File persistence: when requesting an upload URL you pass a persist query parameter.
persist=true— the file is kept indefinitely.persist=false— the file expires after a set period and is automatically deleted.
User Files
User files are personal — each file belongs to a single user and is only accessible by that user. Use this model when the content is user-specific and does not need to be shared.
Required headers
| Header | Description |
|---|---|
x-api-key | Your product API key |
x-user-id | The ID of the user performing the action |
Endpoints
| Method | Path | Description |
|---|---|---|
GET | /v1/signed-url | Get a pre-signed upload URL |
GET | /v1/files | List all files for the user (paginated) |
GET | /v1/files/{fileId} | Retrieve the processed text content of a file |
GET | /v1/files/{fileId}/status | Processing status of a file |
GET | /v1/files/{fileId}/rag | Semantic search within a single file |
DELETE | /v1/files/{fileId} | Delete a file |
Basic use-case flow
Below is the minimal sequence of API calls to upload a file and use it for RAG.
Step 1 — Request an upload URL
────────────────────────────────
GET /v1/signed-url?filename=report.pdf&persist=true
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
Response:
{
"url": "https://s3.amazonaws.com/...", ← pre-signed PUT URL
"file_id": "3f2e1d...", ← keep this ID
"expires_in": 600,
"file_limits": {
"max_size": 10485760,
"allowed_types": ["application/pdf", "text/plain", ...]
}
}Step 2 — Upload the file directly to S3
────────────────────────────────────────
PUT <url from step 1>
Content-Type: application/pdf ← must match the file type
Body: <binary file content>
Response: 200 OK (from S3, no body)Step 3 — Check processing status
───────────────────────────────────────────
GET /v1/files/{fileId}/status
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
Response:
{
"processingStatus": "PROCESSED", ← poll until this is "PROCESSED" or "FAILED"
"steps": [
{ "stepName": "FirstStep", "status": "SUCCEEDED" },
{ "stepName": "ProcessorSelector", "status": "SUCCEEDED" },
{ "stepName": "PdfProcessor", "status": "SUCCEEDED" },
{ "stepName": "ChunkFile", "status": "SUCCEEDED" }
],
"warnings": { "length": false, "empty": false }
}Step 4 — Query the file with RAG
──────────────────────────────────
GET /v1/files/{fileId}/rag?query=What+are+the+main+findings&limit=5
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
Response:
[
{
"content": "The analysis shows ...",
"metadata": { "chunk_id": "...", "file_id": "..." },
"score": 0.94
},
...
]Step 5 (optional) — Retrieve the full processed content
────────────────────────────────────────────────────────
GET /v1/files/{fileId}
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
Response: text/plain — the extracted and processed text of the fileStep 6 (optional) — Delete the file
─────────────────────────────────────
DELETE /v1/files/{fileId}
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
Response: { "message": "File <fileId> deleted successfully." }Tip: if you uploaded with
persist=false, the file will be automatically cleaned up after it expires — you do not need to call DELETE explicitly.
Space Files
Space files belong to a Space — a tenant-scoped container that groups files under a shared context. Spaces have a visibility setting that controls who can access them:
PRIVATE(default) — only the space owner can read, upload, or delete files in the space.PUBLIC— any user within the same tenant can access the space and its files.
Use this model when you want to perform RAG queries over a collection of documents, or when content needs to be shared across users within a tenant.
A Space must exist before files can be uploaded to it. Refer to the Spaces endpoints (POST /v1/spaces) to create one.
Required headers
| Header | Description |
|---|---|
x-api-key | Your product API key |
x-user-id | The ID of the user performing the action |
x-tenant-id | The ID of the tenant the space belongs to |
Endpoints
| Method | Path | Description |
|---|---|---|
GET | /v1/spaces/{spaceId}/files/upload-url | Get a pre-signed upload URL for the space |
GET | /v1/spaces/{spaceId}/files | List all files in the space (paginated) |
GET | /v1/spaces/{spaceId}/files/{fileId} | Retrieve the processed text content of a file |
GET | /v1/spaces/{spaceId}/files/{fileId}/status | Poll the processing status of a file |
GET | /v1/spaces/{spaceId}/rag | Semantic search across all files in the space |
DELETE | /v1/spaces/{spaceId}/files/{fileId} | Delete a file from the space |
Basic use-case flow
Step 0 — Make sure you have a Space
──────────────────────────────────────
POST /v1/spaces
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
x-tenant-id: <tenant-id>
Body:
{
"name": "My Knowledge Base",
"visibility": "PRIVATE" // default; use "PUBLIC" to share with all users in the tenant
}
Response:
{
"id": "space-uuid",
"name": "My Knowledge Base",
...
}Step 1 — Request an upload URL for the space
─────────────────────────────────────────────
GET /v1/spaces/{spaceId}/files/upload-url?filename=guidelines.pdf&persist=true
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
x-tenant-id: <tenant-id>
Response:
{
"url": "https://s3.amazonaws.com/...",
"file_id": "8a1c3f...",
"expires_in": 600,
"file_limits": {
"max_size": 10485760,
"allowed_types": ["application/pdf", "text/plain", ...]
}
}Step 2 — Upload the file directly to S3
────────────────────────────────────────
PUT <url from step 1>
Content-Type: application/pdf
Body: <binary file content>
Response: 200 OK (from S3, no body)Step 3 — Check processing status
───────────────────────────────────────────
GET /v1/spaces/{spaceId}/files/{fileId}/status
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
x-tenant-id: <tenant-id>
Response:
{
"processingStatus": "PROCESSED",
"steps": [
{ "stepName": "FirstStep", "status": "SUCCEEDED" },
{ "stepName": "ProcessorSelector", "status": "SUCCEEDED" },
{ "stepName": "PdfProcessor", "status": "SUCCEEDED" },
{ "stepName": "ChunkFile", "status": "SUCCEEDED" }
],
"warnings": { "length": false, "empty": false }
}Step 4 — Query the entire space with RAG
──────────────────────────────────────────
GET /v1/spaces/{spaceId}/rag?query=What+are+the+holiday+policies&limit=5
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
x-tenant-id: <tenant-id>
Response:
[
{
"content": "Employees are entitled to ...",
"metadata": {
"chunk_id": "...",
"file_id": "...",
"space_id": "..."
},
"score": 0.97
},
...
]The Space RAG endpoint searches across all processed files in the space in a single call. You do not need to query each file individually.
Step 5 (optional) — List files in the space
─────────────────────────────────────────────
GET /v1/spaces/{spaceId}/files?page=1&limit=20&sortBy=createdAt&sortOrder=desc
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
x-tenant-id: <tenant-id>
Response:
{
"data": [
{
"id": "8a1c3f...",
"originalName": "guidelines.pdf",
"processingStatus": "PROCESSED",
"size": 204800,
"createdAt": "2026-03-19T10:00:00Z",
...
}
],
"total": 1,
"page": 1,
"limit": 20
}Step 6 (optional) — Delete a file from the space
──────────────────────────────────────────────────
DELETE /v1/spaces/{spaceId}/files/{fileId}
Headers:
x-api-key: <your-api-key>
x-user-id: <user-id>
x-tenant-id: <tenant-id>
Response: { "message": "File deleted successfully." }Choosing the Right Model
| Scenario | Recommended model |
|---|---|
| A user uploads their own private documents | User Files |
| Multiple documents are shared across a team / tenant | Space Files |
| You need RAG over a single document | User Files — use /v1/files/{fileId}/rag |
| You need RAG over a collection of documents | Space Files — use /v1/spaces/{spaceId}/rag |
| Content is temporary (used once, then discarded) | Either model with persist=false |
| Content is permanent (knowledge base, product data) | Either model with persist=true |
Haufe Files API Docs
Haufe Files API documentation provides detailed information about the API endpoints, request and response formats, and authentication methods. It is designed to help developers integrate with the Haufe Files platform effectively. The whole OpenApi specification can be found [here](/api/openapi). ## API Key Modes This API operates in two distinct modes based on the type of API key used: --- # **Product Mode (General Usage)** Designed for Haufe Products leveraging BYOC (Bring Your Own Content) capabilities to manage customer content through Haufe Files. - Non-Scoped API Keys: - Used by HaufeShop products (e.g., Copilots) - User and tenant management provided by Control Plane - Can access our UI if the IDs match those in Control Plane - Scoped API Keys: - Used by products like HR Assistant or KaaS - Can only access data within their own scope - Automatically provisions tenants and users when needed - User and tenant IDs are prefixed with the scope (e.g., `{scope}:{externalId}`) --- # **Gen AI Mode** A specialized API key type for generative AI tools requiring access to uploaded data for LLM integration. - **Exclusive to HAI (Haufe AI)** - Can read data from all scopes with the appropriate IDs - Limited action set (read-focused capabilities) - Permissions and features will evolve as the product develops --- Feel free to contact the following people in case you have any questions: - Ioana Lefter (Product Owner) - Pavel Khralovich (Tech Lead) - Raul de San Clemente (Software Engineer) - Alejandro Cano (Software Engineer) - Giulia Mainiero (Software Engineer)
Check API health GET
Next Page