Pattern
How to run a nightly backfill without overspending
Pace against a published rate limit, understand what actually costs a credit, and pick refresh candidates so a backfill costs roughly what you budgeted.
A backfill is where enrichment budgets go to die. Every call is billed on success, there is no bulk endpoint, and the naive job (“refresh everything, retry on failure”) multiplies both.
Pick candidates, do not sweep
The first and largest saving is not calling. Refresh what is stale and read:
select company_id, public_id
from companies
where fetched_at < now() - interval '90 days'
and last_read_at > now() - interval '30 days' -- someone actually looks at it
order by last_read_at desc
limit :budget; -- a hard cap, in rows
That limit :budget is the important line. Derive it from money, not from row count:
const MONTHLY_BUDGET_CALLS = 20_000;
const nightlyBudget = Math.floor(MONTHLY_BUDGET_CALLS / 30);
A job that cannot exceed its budget by construction is a job you can leave running.
Pace, do not burst
Your account has a real, published request-per-second allowance, 10 rps by default, and every response tells you where you stand:
X-RateLimit-Limit: 10
X-RateLimit-Remaining: 3
So unlike vendors where 429 is a mystery, you can pace deliberately. Stay under the limit, watch
X-RateLimit-Remaining, and treat a 429 as a bug in your pacing rather than weather.
A small fixed gap plus a concurrency cap is enough:
const CONCURRENCY = 4;
const GAP_MS = 100;
async function paced<T>(items: T[], fn: (item: T) => Promise<void>) {
const queue = [...items];
const workers = Array.from({ length: CONCURRENCY }, async () => {
while (queue.length) {
const item = queue.shift()!;
await fn(item);
await sleep(GAP_MS);
}
});
await Promise.all(workers);
}
Retries that do not double-spend
The failure counts in a badly built backfill are dominated by retries. The same entity attempted four times looks like four failures and can bill more than once. Two rules:
- Retry only
429,502and503. Those are free and transient. - Never retry
400,401,402or404. They will not change on their own, and a retried404costs a credit every single time.
const RETRYABLE = new Set([429, 502, 503]);
if (!res.ok && !RETRYABLE.has(res.status)) {
await recordTerminal(entity, res.status, await res.text());
return; // stop. This will not succeed on attempt two, and a 404 bills again.
}
Log outcomes you can actually read
Recording only success / failed throws away the distinctions you need. Record the status code
and keep the four outcomes separate:
| Outcome | Meaning | Your move |
|---|---|---|
success | Record returned and billed | none |
not_found (404) | Identifier bug, costs a credit | Fix slug extraction, urgently |
rate_limited (429) | Pacing | Lower concurrency |
terminal (401/402) | Credentials or balance | Page someone |
If a 429 outcome is only recorded after three consecutive retries, say so in the metric name.
Otherwise “1.8% rate limited” quietly means “5.4% of calls were throttled”.
Guardrails worth having
- Leave auto-recharge off while a backfill runs. At zero balance calls return
402, which stops the job; with auto-recharge on it just keeps buying credits. - A dry-run mode that counts candidates and prints the projected cost without calling.
- A kill switch the on-call can flip without a deploy.
- Read
X-Credits-Chargedand sum it. It is the only number that is not an estimate.
After the run
Compare fetched_at distribution before and after. If the tail did not move, your candidate query
is selecting the same rows every night, usually because a terminal failure is not being recorded,
so the same unresolvable entities are retried forever.
Next
- Company records: refresh cadences by field.
- Errors: retry semantics in full.
- Billing and credits: exactly what costs a credit.