Documentation menu

Concepts

How files are handled

What happens to a file you send: where it goes, how long it is kept, when it is deleted, and what is logged.

Summary

FileStoredDeleted
Input and output of a direct requestNoThe working copy is deleted when the response has been sent.
Input of a jobYes, until the job finishesWhen the job succeeds, fails or is canceled.
An upload that no job usesYes, for 24 hoursBy an hourly clean-up, once it is 24 hours old.
Output of a jobYes, until it expiresAt expires_at: one hour after the job finishes by default. Earlier if you send DELETE.
Output of a job sent to your own storageNoIt is never written to our storage.

Direct requests

POST /v1/compress, POST /v1/convert and POST /v1/strip-metadata take the file in the request and return the result in the response.

  • The file is passed to a compression engine. It is not written to object storage.
  • The engine works on it in a private directory that belongs to that one request. The directory, with the input and the output in it, is deleted when the response has been sent.
  • The engines have no access to the internet. They can only answer the API.
  • The response carries Cache-Control: no-store.

Job inputs

A job needs its input to outlive one request, so the input is stored.

  • A file sent through an upload is written to private object storage, under your account.
  • A file given as input.url is fetched when the job starts, and copied to the same storage.
  • The input is deleted as soon as the job finishes. This happens on success, on failure and on cancel. The time is recorded and appears in the job's receipt as input_deleted_at.
  • While it works, the engine holds a working copy in a private directory. The directory is removed when the result has been collected, or when the job fails or is canceled. If that removal does not go through, the engine removes the directory about an hour after the work ended.

Job outputs

  • When a job succeeds, its output is written to private object storage. The job's expires_at is set to the time it finished plus the retention period.
  • The retention period is one hour by default. You can change the account default in the dashboard, and set it for one job with retention_seconds. The range is 60 seconds to 24 hours (86,400 seconds).
  • At expires_at the output is deleted by a timer that belongs to the job. It does not wait for a periodic clean-up.
  • Until then the output can be downloaded with result.download_url, or with your API key at GET /v1/jobs/{id}/output. The download link stops working at the same time.
  • After deletion, the job shows files_deleted: true and download_url: null.
keep the result for ten minutes
{
  "operation": "compress",
  "input": { "upload": "upl_Zk3q8WcT1nRb5LxYp0Ha" },
  "retention_seconds": 600
}

Deleting on demand

DELETE /v1/jobs/{id} deletes the job's input and output at once. If the job is still queued or processing, it is canceled first. You do not have to wait for expires_at. See the jobs reference.

Output to your own storage

If you give a job an output.url, the result is sent to that URL with an HTTP PUT and is not written to our storage. The job has no download_url, and nothing of the output remains with us once the job finishes. The input is still stored until the job finishes. See Output to your storage.

What is stored, and for how long

What is kept about a request and its files, files first, then records:

WhatHoldsKept
The file of a direct requestThe file and the result, as a working copy on the engine's own disk.Until the response has been sent.
The input of a jobThe file.Until the job ends: success, failure or cancel.
The output of a jobThe result.For retention_seconds: 1 minute to 24 hours, 1 hour by default. Or until you send DELETE.
An upload that no job usedThe file.24 hours.
An upload in parts that was not completedThe parts sent so far.24 hours.
The record of a jobFile name, options, webhook and output addresses, output headers, metadata, result figures, hashes of the input and the output.30 days after the job ends.
The job listJob id, key id, operation, status, kind of file, error code, whether it was billed, times. No file names.30 days after the job was created.
Idempotency recordsThe Idempotency-Key, a hash of the request body, and the response that was sent.24 hours.
Request recordsRequest id, key id, time, operation, kind of file, sizes, engine time, billing meter and cost, error code. No file names, no contents.400 days.
Daily usage totalsCounts, quantities and amounts per day and meter.While the account exists.
LogsOur own log lines: request ids, error codes, timings. Request URLs are not logged, because a URL can carry a file name.The log provider's retention.
  • The 24-hour and 30-day deletions of uploads, idempotency records, the job list and request records are made by a clean-up that runs once an hour, so they can come up to an hour late. The input and the output of a job are deleted by the job itself, on time.
  • Things you create stay until you delete them: presets (a name and an options object), webhook endpoints (a URL and its signing secret), and API keys (stored as a hash).
  • An upload in parts also leaves a small record: its id, its size, the size and number of its parts, and when it was created and completed. It holds no file name and no content, and it is not removed when the upload is used.

What is recorded

Three kinds of record are kept about a request. None of them holds file contents.

RecordHoldsDoes not hold
Usage log, one row per requestRequest id, account and key ids, time, operation, kind of file, outcome, error code, billing meter and amount, input and output size in bytes, processing time.File names. File contents.
Job list, one row per jobJob id, account and key ids, operation, status, kind of file, error code, whether it was billed, creation and completion times.File names. File contents.
Job recordEverything in the job object, including the file name you gave and your metadata. Also the URLs and headers you supplied for input, output and webhooks.File contents.
  • The job record is deleted 30 days after the job finishes. After that the job and its receipt can no longer be fetched by id. Fetch a receipt you want to keep before then.
  • A row of the job list is deleted 30 days after the job was created.
  • A row of the usage log is deleted 400 days after the request. It is kept that long because it is the basis of your invoices.
  • Platform request logs, which would record the URL of each request, are turned off. A URL can carry a file name in ?filename=.
  • Engine logs record the operation, formats, sizes and processing time of each request. When a tool fails, the log holds the name of the tool, its exit code and the amount of error output it produced. The error output itself is not logged, because it can quote tags and other text from inside the file. The file name you sent is not logged either.
  • An API key is stored as a SHA-256 hash, together with its first characters and its last four characters so that the dashboard can show which key is which. The key itself is not stored.

Receipts

Every finished job has a receipt at GET /v1/jobs/{id}/receipt. For a job that succeeded, it states the SHA-256 hashes and sizes of the input and the output. For every job, it states when the job was created and completed, when the input and the output were deleted, and whether the output went to your own storage. The statement is signed with an Ed25519 key, so you can keep it as evidence and check later that it has not been altered. The public keys are published at GET /v1/receipt_keys. See Receipts.

A deletion time in a receipt is null until that deletion has happened. To get a receipt that shows the output deleted, fetch it after expires_at or after a DELETE.

What is not promised

  • This is not end-to-end encryption. Requests travel over HTTPS, but the service must read a file to compress it. While a file is being processed, and while a job's input or output is stored, it exists on our side in readable form.
  • A download link is a secret. Anyone who has result.download_url can download the output until it expires. Do not log it or pass it on to parties who should not have the file.
  • Deletion covers our copies. A file fetched from your input.url, sent to your output.url, or downloaded by you is yours to manage.
  • Webhook payloads leave our systems. A webhook carries the job object, including the file name and a download link, to the URL you registered.