Concepts
Limits
Rate limits, concurrency limits, file limits for each plan, and the monthly spend cap.
Rate limit
The rate limit applies to direct requests: POST /v1/compress, POST /v1/convert and POST /v1/strip-metadata. It is counted per account, across all of the account's keys of the same mode. Live keys and test keys are counted separately.
It is a token bucket. The bucket holds twice your plan's requests per second. Each request takes one token. The bucket refills at your plan's requests per second. So you can send a burst of twice the rate, and after that the sustained rate.
Example: on a plan with 20 requests per second, the bucket holds 40. You can send 40 requests at once. After that, 20 more are admitted each second. A bucket that is empty is full again after 2 seconds without requests.
A request that arrives when the bucket is empty is refused with status 429 and code rate_limited.
Every other endpoint that needs a key (uploads, creating and reading jobs, presets, usage, webhook endpoints) shares a second, looser bucket, so that watching many jobs does not use up the rate for processing. It refills at ten times your plan's requests per second, with a minimum of 20 a second, and holds twice that. For test keys it is 20 a second on every plan. A request over it gets the same 429 rate_limited with Retry-After, but these responses do not carry the RateLimit-* headers.
Upload allowance
Uploads cost nothing, and an upload no job uses is kept for a day. To stop that being used as free storage, an account can upload a limited number of bytes in any 24 hours: twenty times its plan's largest job file (for example 100 GB on Growth), or 500 MB for test keys. Every upload counts, whether or not a job used it. Over the allowance, an upload is refused with 429, upload_allowance_reached. Job inputs given as a URL do not count.
Rate limit headers
Once the API key has been accepted, every response from a direct endpoint carries these headers, on success and on error. A response that fails authentication does not carry them.
| Header | Meaning |
|---|---|
RateLimit-Limit | The size of the bucket: twice your requests per second. |
RateLimit-Remaining | Whole tokens left after this request. |
RateLimit-Reset | Seconds until the bucket is full again. On a refused request, seconds until a request can be admitted. |
Retry-After | Sent with 429 and with 503 engine_busy. Seconds to wait before the next attempt. |
HTTP/1.1 429 Too Many Requests
Content-Type: application/json; charset=utf-8
RateLimit-Limit: 40
RateLimit-Remaining: 0
RateLimit-Reset: 1
Retry-After: 1
Smol-Request-Id: req_Qm4tY8sLx02BnVc7JdPe
Smol-Version: 2026-10-01
{
"error": {
"type": "rate_limit_error",
"code": "rate_limited",
"message": "Too many requests. Slow down and retry.",
"param": null,
"request_id": "req_Qm4tY8sLx02BnVc7JdPe",
"doc_url": "https://smolmac.com/docs/reference/errors#rate_limited"
}
}See Errors and retries for a retry loop that respects Retry-After.
Concurrency limits
There are two, both counted per account.
- Concurrent direct requests. The number of direct requests being processed at once. One over the limit is refused with
429, concurrency_limited andRetry-After: 1. This check comes before the rate limit, so a request refused here does not take a token. - Concurrent jobs. The number of jobs that are queued or processing. Creating one more with
POST /v1/jobsis refused with429,concurrency_limitedandRetry-After: 5. A job stops counting when it succeeds, fails or is canceled.
Rates by plan
| Plan | Requests per second | Burst | Concurrent direct requests | Concurrent jobs |
|---|---|---|---|---|
| No plan (test keys only) | 2 | 4 | 2 | 1 |
| Pay as you go | 5 | 10 | 5 | 3 |
| Starter | 20 | 40 | 20 | 10 |
| Growth | 50 | 100 | 50 | 30 |
| Scale | 150 | 300 | 150 | 100 |
| Enterprise | 500 | 1,000 | 500 | 500 |
File limits by plan
| Plan | Direct file | Job file | Video or audio length | Pixels | PDF pages |
|---|---|---|---|---|---|
| No plan (test keys only) | 25 MB | 100 MB | 1 minute | 50 megapixels | 200 |
| Pay as you go | 25 MB | 500 MB | 10 minutes | 100 megapixels | 1,000 |
| Starter | 25 MB | 2 GB | 1 hour | 100 megapixels | 2,000 |
| Growth | 25 MB | 5 GB | 3 hours | 100 megapixels | 2,000 |
| Scale | 25 MB | 5 GB | 3 hours | 100 megapixels | 2,000 |
| Enterprise | 25 MB | 5 GB | 3 hours | 100 megapixels | 2,000 |
- Direct file is the largest file a direct request accepts. A larger one is refused with
413, file_too_large. Use a job. - Job file is the largest input a job accepts, whether it comes from an upload or from
input.url. An upload of more than 95 MB is sent in parts. - Pixels is width times height, for an image and for a video frame.
- A file over the length, pixel or page limit is refused with
422, limit_exceeded. The message states the file's value and the limit.
MB and GB here are binary: 25 MB is 26,214,400 bytes.
Limits that do not depend on the plan
| Limit | Value | When it is exceeded |
|---|---|---|
| Frames in an animated image | 2,000 | 422 limit_exceeded |
| Processing time of a direct request | 120 seconds | 504 timeout. Use a job. |
| Processing time of a job | 4 hours | The job fails with code timeout. |
| Time a job waits for capacity | 30 minutes | The job fails with code timeout. It is not billed. |
| Upload content sent in one request | 95 MB | A larger upload is sent in parts of 64 MB. Sending more than 95 MB in one request returns 413 file_too_large. |
| Time an unused upload is kept | 24 hours | The upload is deleted. |
| Presets per account | 100 | 409 preset_limit_reached |
| Days covered by one usage query | 366 | 400 bad_request |
| JSON request body | 64 KB | 400 bad_request |
| Webhook endpoints per account | 10 | 400 bad_request |
| Active API keys per account | 20 | The dashboard refuses to create another until one is revoked. |
| Addresses or ranges in the allow-list of a key | 20 | The dashboard refuses the list. |
Length of an Idempotency-Key | 1 to 255 characters | 400 bad_request |
Headers in a job's output.headers | 20, with names up to 100 characters and values up to 2,000 | 400 bad_request |
| Redirects followed when a job fetches input.url | 3 | The job fails with code input_fetch_failed. |
Entries in a job's metadata | 20, with keys up to 40 characters and values up to 500 | 400 bad_request |
image.resize width and height | 100 to 9,999 pixels | Values outside the range are moved to the nearest end of it. |
Test keys
A key starting smol_test_ is limited to 2 requests per second, with a burst of 4, on every plan. A direct request with a test key accepts files up to 10 MB, and a job with a test key accepts an upload of up to 25 MB. A test key cannot start an upload in parts. The concurrency limits and the pixel, page and length limits are those of the account's plan. See Test mode.
Monthly spend cap
Every account has a spend cap. It limits the value of billable usage in one calendar month, counted in UTC. Usage is counted at the unit prices, before any usage included in a plan is taken off. Plan fees are not counted.
| Plan | Default cap |
|---|---|
| Pay as you go | $100 |
| Starter | $250 |
| Growth | $1,000 |
| Scale | $5,000 |
| Enterprise | $50,000 |
You can set your own cap in the dashboard. Once the month's usage reaches the cap:
- Direct requests and new jobs with a live key are refused with
402, spend_cap_reached. - Requests and jobs that are already running finish, and are billed as usual.
- Work that is still running counts towards the cap at the most it could cost, until it finishes and its real price is known. So a burst of long jobs started together cannot run the month far past the cap: once what has been spent plus what is running could reach it, new work is refused until some of it finishes. One request is always let through while the month is under the cap.
- Test keys keep working. Reading jobs and downloading results keeps working.
The cap is checked when a request starts, not when it ends. The month's usage can therefore finish a little above the cap, by at most the cost of one request. The count returns to zero at the start of the next month.