Guides
Output to your storage
Have a job upload its result straight to your own bucket through a presigned URL, so that Smol never stores the output.
How it works
By default a job stores its result for a short time and gives you a link to download it. With output.url the job sends the result to a URL you supply instead, with one HTTP PUT. The usual target is a presigned upload URL for an object in your own bucket.
- Your server asks your storage provider for a presigned PUT URL for the object the result should become.
- You create a job with that URL in
output.url. - When processing is done, the result is streamed to the URL.
- The job succeeds only if your storage accepted the upload.
This saves you the download and the re-upload, and it means the output is never written to our storage. Output delivery is a feature of jobs. Direct requests return the file in the response and have no output option.
Create a presigned URL
A presigned URL is a normal https URL with a signature in its query string. It lets whoever holds it upload one object, to one key, until it expires, without holding your storage credentials. Amazon S3, Cloudflare R2, Google Cloud Storage, Backblaze B2 and other S3-compatible services can all issue one.
import { PutObjectCommand, S3Client } from "@aws-sdk/client-s3";
import { getSignedUrl } from "@aws-sdk/s3-request-presigner";
const s3 = new S3Client({ region: "eu-west-1" });
// Do not set ContentType here: Smol sets the Content-Type of the result itself.
const outputUrl = await getSignedUrl(
s3,
new PutObjectCommand({ Bucket: "my-bucket", Key: "compressed/report.pdf" }),
{ expiresIn: 6 * 60 * 60 },
);import boto3
s3 = boto3.client("s3", region_name="eu-west-1")
# Do not set ContentType here: Smol sets the Content-Type of the result itself.
output_url = s3.generate_presigned_url(
"put_object",
Params={"Bucket": "my-bucket", "Key": "compressed/report.pdf"},
ExpiresIn=6 * 60 * 60,
)Two things to get right when you create the URL:
- Expiry. The URL is used at the end of the job, not when the job is created. It must still be valid then. Allow for time in the queue and for processing: an hour is enough for images and documents, and a long video needs more.
- Content type. If the signature covers a content type, the upload must send exactly that type or the storage provider rejects it. The upload is sent with the content type of the result, for example
application/pdforvideo/mp4. The simplest approach is not to sign a content type at all.
Create the job
curl -X POST "https://api.smolmac.com/v1/jobs" \
-H "Authorization: Bearer $SMOL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"operation": "compress",
"input": { "upload": "upl_Zk3Vb9QeT1mXc7HsW0yN" },
"options": { "pdf": { "quality": "small" } },
"output": {
"url": "https://my-bucket.s3.eu-west-1.amazonaws.com/compressed/report.pdf?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Signature=..."
}
}'| Field | Type | Description |
|---|---|---|
| output.url * | string | Where to PUT the result. It must use https on the default port, with a public host name and no user name or password in the URL. |
| output.headers | object | Extra request headers to send with the upload, as strings. Up to 20 headers, names of up to 100 characters, values of up to 2,000 characters. |
A URL that does not meet the rules is refused when the job is created, with 400 bad_request and param set to output.url. The input side is unchanged: it can be an upload or an input.url. With a presigned GET URL as the input and a presigned PUT URL as the output, a file goes from your bucket, through the job, and back to your bucket.
The upload request
This is the request your storage receives:
PUT /compressed/report.pdf?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Signature=... HTTP/1.1
Host: my-bucket.s3.eu-west-1.amazonaws.com
Content-Type: application/pdf
Content-Length: 9627345
<the result file>Content-TypeandContent-Lengthare always set by the job. Putting either inoutput.headersis refused with400 bad_request, as areHost,Cookie, the headers that describe the connection, and any name or value that is not a plain header.- The headers in
output.headersare added as they are. Use them for anything your storage requires on the request. For example, an Azure Blob Storage SAS URL needsx-ms-blob-type: BlockBlob, and an S3 URL that was signed with server-side encryption needs the matchingx-amz-server-side-encryptionheader. - Any
2xxanswer counts as success. Redirects are not followed, so the URL must be the final address.
"output": {
"url": "https://myaccount.blob.core.windows.net/compressed/report.pdf?sv=2024-11-04&sig=...",
"headers": { "x-ms-blob-type": "BlockBlob" }
}The finished job
The job reaches succeeded after your storage has accepted the upload. The result object describes the file as usual, but there is nothing to download from us:
{
"id": "job_8Hq2LmZx4TnV0cRb7KpW",
"object": "job",
"status": "succeeded",
"result": {
"kind": "pdf",
"output_format": "pdf",
"original_size": 64182301,
"output_size": 9627345,
"savings_percent": 85.0,
"kept_original": false,
"download_url": null,
"download_expires_at": null
},
"billed": true,
"files_deleted": true
}download_url is null, and files_deleted is true as soon as the job is done. Webhooks work as with any job: the job.succeeded event tells you the object is in your bucket.
When the upload fails
If your storage does not answer the PUT with a 2xx status, or cannot be reached, the job fails. The upload is attempted once.
{
"id": "job_8Hq2LmZx4TnV0cRb7KpW",
"object": "job",
"status": "failed",
"result": null,
"error": {
"type": "invalid_request_error",
"code": "output_upload_failed",
"message": "The output could not be uploaded to the output URL (HTTP 403)."
},
"billed": false,
"files_deleted": true
}The job is not billed. The processed file is discarded: no copy is kept on our side to retry from or to download. To try again, create a new presigned URL and a new job. The status in the message is the one your storage returned. The common causes:
| Status from your storage | Likely cause |
|---|---|
| 403 | The URL expired before the job finished, the signature covers a different content type, or a required header is missing. |
| 400 | A header your storage requires on the request was not sent. Add it to output.headers. |
| 301, 302, 307 | The URL redirects, often because of the wrong bucket region. Use the final address. |
| No status in the message | The host could not be reached. |
What Smol keeps
| File | With output.url |
|---|---|
| Input | Deleted as soon as the job reaches a final status, as with every job. |
| Output | Never written to our storage. It is streamed from the processing engine to your URL and is gone from the engine when the job ends. |
retention_seconds has no effect on a job with output.url, because there is no stored output to retain. The record of the job, without files, stays readable for 30 days. That record holds the output.url and output.headers you sent, although the job object does not show them. Give a presigned URL an expiry that is no longer than the job needs, and do not put long-lived credentials in the headers.
The receipt of such a job states output_delivered_to_customer_storage: true. The field says the job had an output.url. Read it together with status: only a job that succeeded delivered its output. For the whole picture, see How files are handled.