Blog
Build your own profile scraper or call an enrichment API: the cost nobody prices in
A profile scraper is cheap to write and costly to keep alive. Here is the recurring cost you do not see until the markup changes and the run returns nothing.
A scraper that reads profile and company pages takes an afternoon to write and a year to keep alive. The afternoon is the part you price. The year is the part that decides whether this was the right call.
The argument in one paragraph
A scraper has no fixed cost and a large variable one. The page markup changes and your selectors return empty strings. An IP gets blocked and you add a proxy pool. The proxy pool gets blocked and you rotate it. Pacing drifts and a whole run comes back as challenge pages instead of data. None of that work ships a feature. An enrichment API moves that variable cost off your roadmap and turns it into a line item: one credit per answered lookup, billed on success, with the retrieval and the blocking someone else’s problem.
What you own the day you write the scraper
Four jobs, and all four recur forever.
- Parsers that track a page you do not control. A profile page is HTML built for browsers, not
for you. When the class names change, your
current_titleselector returns""and nothing errors. You find out from a customer, not a stack trace. - Proxies and the block that follows them. One IP reading thousands of pages gets rate limited, then blocked, then served a login wall. You buy a proxy pool to spread the load, and now you maintain the pool.
- Pacing you have to guess. There is no published limit to code against, so you tune request spacing by watching for the point where responses turn into challenge pages, and you retune it every time the defenses change.
- Provenance you have to build. When someone asks in six months where a field came from and how old it is, a scraper gives you a row with no answer attached. You have to design and store that yourself, or you cannot answer.
What the same work looks like as a call
The enrichment side is one request with an identifier. Auth is an ApiKey header, the base is
https://api.triguna.ai/v1, and the person endpoint takes a slug in profile_id:
const res = await fetch(
'https://api.triguna.ai/v1/people/profile?profile_id=williamhgates',
{ headers: { ApiKey: process.env.TRIGUNA_API_KEY } },
);
const profile = await res.json();
import os, requests
res = requests.get(
'https://api.triguna.ai/v1/people/profile',
params={'profile_id': 'williamhgates'},
headers={'ApiKey': os.environ['TRIGUNA_API_KEY']},
)
profile = res.json()
The same shape enriches a company through GET /v1/companies/details, keyed on the URL slug rather
than the display name. Two endpoints, one identifier each. See the Profile API
and the Company API for the response field by field.
The provenance you stop building
The four headers a scraper never gives you arrive on every response, so you can tell where a value came from and how stale it is before you trust it:
X-Data-Source: profile
X-Fetched-At: 2026-09-11T14:02:55Z
X-Data-Age-Seconds: 43
X-Credits-Charged: 1
Provenance lives in the headers, never in the response body. That is the difference between “here is the exact state we received, on this date” and a bare row you have to defend from memory. The One structured log line per enrichment call post shows how to keep that record cheaply, and Store the raw payload explains why a stored response outlives the service that produced it.
The failure you do not retry
A scraper fails silently: a selector goes stale and you keep writing empty strings to the database. A
call fails loudly, with a status and a header that tells you what to do. A rate limit is a 429 with
a wait attached:
HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Remaining: 0
A 429 is free and you retry it after the wait. A 404 is a real lookup that costs one credit and
you never retry it, because a retried 404 costs a credit every time. Knowing which is which is a
contract, not a guess. Which enrichment errors to retry
draws the full line.
Where buying does not help
Be honest about the boundary, because it decides the answer. An enrichment API answers “given this identifier, return this record.” It is N calls for N entities, with no bulk path and no search endpoint, so it does not find identifiers you do not already have. If your actual job is discovery, crawling a list of people by some filter, an enrichment API is the wrong tool and no amount of it substitutes for a source of identifiers. If your job is turning identifiers you already hold into current, attributable records, the build side is pure cost with no upside.
Running the enriched lookups at volume is its own small discipline: pace them under the published
limit and checkpoint so a restart does not repeat paid calls. The nightly backfill
guide covers that, and Pricing lists what each status code costs, including the 404.
The rule
Write the scraper when the data is the product and you have the team to defend it forever. Call the API when the data is an input and you would rather ship. Most of the time the data is an input. The team behind Triguna built the second thing because they kept paying for the first.
Next
- Profile API: the
GET /v1/people/profileresponse, field by field. - Company API:
GET /v1/companies/detailsand the slug it keys on. - Migrating off Proxycurl: moving profile and company enrichment onto two endpoints, with before and after code.
- Handling partial records: what to do when a field you expected comes back null.
- Pricing: what each status code costs, including the 404 you paid for.