Blog
Field notes on enrichment
Mostly things we got wrong first. Outcome distributions from real traffic, rate limits that ignore your request rate, and what happens when a provider goes dark. Subscribe by RSS.
Latest
Apollo.io vs a two-endpoint enrichment API: when a GTM platform fits and when a keyed lookup does
Apollo.io matches partial fields against a 240M-contact GTM database. A two-endpoint enrichment API answers a keyed lookup with provenance in the headers.
All posts
- Clay vs a direct enrichment API: when a data marketplace and a single contract each fit Clay is a data marketplace that runs waterfalls across many providers. A direct enrichment API is one documented contract. Here is when each one fits.
- Clearbit's standalone API is deprecated: what a two-endpoint enrichment API replaces Clearbit's standalone API was deprecated after the HubSpot acquisition. Here is what a two-endpoint enrichment API covers, and where it stops.
- ScrapIn vs an on-demand enrichment API: real-time scraping or a two-endpoint contract ScrapIn is a real-time scraping layer with a broad surface. An on-demand enrichment API does two endpoints with provenance in the headers. When each fits.
- Bright Data for LinkedIn data versus a two-endpoint enrichment API, and where each one fits Bright Data sells LinkedIn datasets and a scraper API billed per record. Here is how that differs from on-demand enrichment with provenance on every call.
- Apify LinkedIn scraper actors versus a managed enrichment API, and the axes that actually differ Apify sells LinkedIn scraper actors billed per result. Here is how that model differs from a managed enrichment API with provenance and a fixed lookup contract.
- Coresignal vs People Data Labs: two dataset providers, and when on-demand enrichment fits better Coresignal and People Data Labs are both stored-dataset providers. Here is how the two differ, and when on-demand enrichment fits your use case better.
- Build your own profile scraper or call an enrichment API: the cost nobody prices in A profile scraper is cheap to write and costly to keep alive. Here is the recurring cost you do not see until the markup changes and the run returns nothing.
- Check your credit balance before a batch run, because a 402 halfway through stalls everything One credit per lookup, and a 404 bills too, so a long run drains credits faster than it succeeds. Read the balance first so a 402 never stops you midway.
- One structured log line per enrichment call, so the bill and the bug stay answerable weeks later Every enrichment response carries its cost and age in the headers, then it is gone. Log one line per call so the bill, a failure, and freshness stay answerable.
- No bulk endpoint: enriching a list is N calls, paced with bounded concurrency under the limit There is no bulk endpoint, so enriching N records means N calls. Here is the bounded concurrency pattern that stays under the 10 rps limit and collects errors.
- A fresh read and a cache hit cost the same credit: when to send use_cache=false A cache hit and a forced fresh read both bill one credit, so cost is never the reason to send use_cache=false. Here is when a source read is worth it.
- Key your enriched records on the entity urn, because the slug you looked them up by will change The slug you enrich by is a vanity URL the owner can rename. Key your records on the entity urn instead, or a renamed profile quietly becomes a second row.
- Test your enrichment integration on the five free credits, before you wire up billing An auth failure, a rate limit, and a malformed identifier all cost zero credits. So you can exercise almost your whole integration on the five signup credits.
- Provenance lives in the headers: how to tell where enrichment data came from and how old it is Provenance on an enrichment response lives in the headers, not the body. Here is what the six response headers tell you and the decision each one should drive.
- Which enrichment API errors to retry, and which ones cost a credit if you do Only 429, 502, and 503 are worth retrying. Retry a 404 and you pay a credit every time. The full status table, which codes are free, and the backoff that works.
- How often should you re-enrich a record? Set the cadence by field, not by calendar Contact details, job titles, and firmographics decay at very different rates. A refresh window set per field keeps data fresh without re-fetching what has not changed.
- Proxycurl alternatives: how the replacements actually differ, and where Triguna fits Proxycurl shut down and its users need a replacement. Here is how the scrapers and dataset providers differ, and where a two-endpoint enrichment API fits.
- What an enrichment run costs, and how to reconcile the bill against the headers One credit per answered lookup, no bulk endpoint, and 404s still bill. Here is how to estimate an enrichment run and reconcile it against X-Credits-Charged.
- Why profile and company lookups return 404 when the identifier looks correct A display name, a urn:li: prefix, or the wrong url field 404s an enrichment lookup, and every 404 still costs a credit. Fix the four identifier traps first.
- Migrating off Proxycurl: mapping profile and company enrichment to two endpoints Proxycurl shut down on 4 July 2025 after LinkedIn sued. Move your profile and company enrichment onto two endpoints, with before and after code.
- Why a cache hit still costs a credit, and what the cache contract actually gives you A cache hit on an enrichment API costs the same credit as a fresh fetch. Here is why, what the 24 hour freshness window covers, and how to store on top.
- Store the raw payload, not just the fields you parse today The cheapest decision in any enrichment integration is keeping the whole response. It costs bytes, it survives vendors, and it means you never pay for the same call twice.
- When your enrichment provider goes dark A vendor's failure rate went 0.3%, 7.1%, 45.6%, 82.9%, 98.1%, then 100% over six weeks. We read it as our own bug the whole time. What we should have been watching.
- Rate limits that have nothing to do with your request rate Our heaviest day ran 19,154 calls and was never throttled. A day running 201 calls was throttled 64% of the time. Why tuning a throttle against 429 is usually wasted work.
- What 130,220 enrichment API calls taught us Seven weeks of logged calls against a third-party enrichment API, and what the outcome distribution actually said about data quality, retries, and instrumentation.
Built from these lessons
A stable response shape, a stored copy behind every lookup, and provenance headers on every response.