Concepts
How files are handled
What happens to a file you send: where it goes, how long it is kept, when it is deleted, and what is logged.
Summary
| File | Stored | Deleted |
|---|---|---|
| Input and output of a direct request | No | The working copy is deleted when the response has been sent. |
| Input of a job | Yes, until the job finishes | When the job succeeds, fails or is canceled. |
| An upload that no job uses | Yes, for 24 hours | By an hourly clean-up, once it is 24 hours old. |
| Output of a job | Yes, until it expires | At expires_at: one hour after the job finishes by default. Earlier if you send DELETE. |
| Output of a job sent to your own storage | No | It is never written to our storage. |
Direct requests
POST /v1/compress, POST /v1/convert and POST /v1/strip-metadata take the file in the request and return the result in the response.
- The file is passed to a compression engine. It is not written to object storage.
- The engine works on it in a private directory that belongs to that one request. The directory, with the input and the output in it, is deleted when the response has been sent.
- The engines have no access to the internet. They can only answer the API.
- The response carries
Cache-Control: no-store.
Job inputs
A job needs its input to outlive one request, so the input is stored.
- A file sent through an upload is written to private object storage, under your account.
- A file given as
input.urlis fetched when the job starts, and copied to the same storage. - The input is deleted as soon as the job finishes. This happens on success, on failure and on cancel. The time is recorded and appears in the job's receipt as
input_deleted_at. - While it works, the engine holds a working copy in a private directory. The directory is removed when the result has been collected, or when the job fails or is canceled. If that removal does not go through, the engine removes the directory about an hour after the work ended.
Job outputs
- When a job succeeds, its output is written to private object storage. The job's
expires_atis set to the time it finished plus the retention period. - The retention period is one hour by default. You can change the account default in the dashboard, and set it for one job with
retention_seconds. The range is 60 seconds to 24 hours (86,400 seconds). - At
expires_atthe output is deleted by a timer that belongs to the job. It does not wait for a periodic clean-up. - Until then the output can be downloaded with
result.download_url, or with your API key atGET /v1/jobs/{id}/output. The download link stops working at the same time. - After deletion, the job shows
files_deleted: trueanddownload_url: null.
{
"operation": "compress",
"input": { "upload": "upl_Zk3q8WcT1nRb5LxYp0Ha" },
"retention_seconds": 600
}Deleting on demand
DELETE /v1/jobs/{id} deletes the job's input and output at once. If the job is still queued or processing, it is canceled first. You do not have to wait for expires_at. See the jobs reference.
Output to your own storage
If you give a job an output.url, the result is sent to that URL with an HTTP PUT and is not written to our storage. The job has no download_url, and nothing of the output remains with us once the job finishes. The input is still stored until the job finishes. See Output to your storage.
What is stored, and for how long
What is kept about a request and its files, files first, then records:
| What | Holds | Kept |
|---|---|---|
| The file of a direct request | The file and the result, as a working copy on the engine's own disk. | Until the response has been sent. |
| The input of a job | The file. | Until the job ends: success, failure or cancel. |
| The output of a job | The result. | For retention_seconds: 1 minute to 24 hours, 1 hour by default. Or until you send DELETE. |
| An upload that no job used | The file. | 24 hours. |
| An upload in parts that was not completed | The parts sent so far. | 24 hours. |
| The record of a job | File name, options, webhook and output addresses, output headers, metadata, result figures, hashes of the input and the output. | 30 days after the job ends. |
| The job list | Job id, key id, operation, status, kind of file, error code, whether it was billed, times. No file names. | 30 days after the job was created. |
| Idempotency records | The Idempotency-Key, a hash of the request body, and the response that was sent. | 24 hours. |
| Request records | Request id, key id, time, operation, kind of file, sizes, engine time, billing meter and cost, error code. No file names, no contents. | 400 days. |
| Daily usage totals | Counts, quantities and amounts per day and meter. | While the account exists. |
| Logs | Our own log lines: request ids, error codes, timings. Request URLs are not logged, because a URL can carry a file name. | The log provider's retention. |
- The 24-hour and 30-day deletions of uploads, idempotency records, the job list and request records are made by a clean-up that runs once an hour, so they can come up to an hour late. The input and the output of a job are deleted by the job itself, on time.
- Things you create stay until you delete them: presets (a name and an options object), webhook endpoints (a URL and its signing secret), and API keys (stored as a hash).
- An upload in parts also leaves a small record: its id, its size, the size and number of its parts, and when it was created and completed. It holds no file name and no content, and it is not removed when the upload is used.
What is recorded
Three kinds of record are kept about a request. None of them holds file contents.
| Record | Holds | Does not hold |
|---|---|---|
| Usage log, one row per request | Request id, account and key ids, time, operation, kind of file, outcome, error code, billing meter and amount, input and output size in bytes, processing time. | File names. File contents. |
| Job list, one row per job | Job id, account and key ids, operation, status, kind of file, error code, whether it was billed, creation and completion times. | File names. File contents. |
| Job record | Everything in the job object, including the file name you gave and your metadata. Also the URLs and headers you supplied for input, output and webhooks. | File contents. |
- The job record is deleted 30 days after the job finishes. After that the job and its receipt can no longer be fetched by id. Fetch a receipt you want to keep before then.
- A row of the job list is deleted 30 days after the job was created.
- A row of the usage log is deleted 400 days after the request. It is kept that long because it is the basis of your invoices.
- Platform request logs, which would record the URL of each request, are turned off. A URL can carry a file name in
?filename=. - Engine logs record the operation, formats, sizes and processing time of each request. When a tool fails, the log holds the name of the tool, its exit code and the amount of error output it produced. The error output itself is not logged, because it can quote tags and other text from inside the file. The file name you sent is not logged either.
- An API key is stored as a SHA-256 hash, together with its first characters and its last four characters so that the dashboard can show which key is which. The key itself is not stored.
Receipts
Every finished job has a receipt at GET /v1/jobs/{id}/receipt. For a job that succeeded, it states the SHA-256 hashes and sizes of the input and the output. For every job, it states when the job was created and completed, when the input and the output were deleted, and whether the output went to your own storage. The statement is signed with an Ed25519 key, so you can keep it as evidence and check later that it has not been altered. The public keys are published at GET /v1/receipt_keys. See Receipts.
A deletion time in a receipt is null until that deletion has happened. To get a receipt that shows the output deleted, fetch it after expires_at or after a DELETE.
What is not promised
- This is not end-to-end encryption. Requests travel over HTTPS, but the service must read a file to compress it. While a file is being processed, and while a job's input or output is stored, it exists on our side in readable form.
- A download link is a secret. Anyone who has
result.download_urlcan download the output until it expires. Do not log it or pass it on to parties who should not have the file. - Deletion covers our copies. A file fetched from your
input.url, sent to youroutput.url, or downloaded by you is yours to manage. - Webhook payloads leave our systems. A webhook carries the job object, including the file name and a download link, to the URL you registered.