Blog
How often should you re-enrich a record? Set the cadence by field, not by calendar
Contact details, job titles, and firmographics decay at very different rates. A refresh window set per field keeps data fresh without re-fetching what has not changed.
The usual answer to “how often should we refresh our enrichment data” is a number: every quarter, every ninety days, once a year. A single number is the wrong shape for the problem, because the fields inside one record do not go stale together.
One record, several clocks
Published decay figures for B2B contact data land somewhere between twenty and seventy percent per year depending on the sector, but the aggregate hides the thing you can actually act on. A person’s email survives a job change badly. Their title survives it not at all. The company’s headquarters, industry, and legal name barely move for years.
So a record is not one clock. It is three, running at very different speeds:
- Volatile. Job title, seniority, current employer, direct email. These break on a single career move, and a career move leaves no signal in your database until you look.
- Slow. Headcount band, funding stage, tech stack, location. These drift over quarters, not weeks.
- Effectively static. Founding year, company domain, canonical name. If these change you have a merger or a rebrand, which is an event you hear about, not a cadence you schedule.
A calendar refresh either re-fetches the static fields far too often, or lets the volatile ones rot for a quarter. It cannot do both, because it treats one record as one clock.
Every refresh is a billed call, so target it
The reason this matters is cost, not tidiness. A refresh is a fresh lookup, and a fresh lookup is billed on success like any other call. Re-fetching a million records every ninety days because the emails might have moved means paying for a million founding years that did not.
The lever is that the Profile API and the Company API are separate calls against separate identifiers. Person volatility and company volatility are on different clocks, so refresh them on different schedules. Re-enrich the person when the title is what you need current; re-enrich the company far less often, because the firmographics you read from it barely moved.
Track staleness per field group, not per row
You cannot refresh by field group if you only record one fetched_at for the whole row. Give each
clock its own timestamp:
alter table people
add column contact_fetched_at timestamptz, -- volatile: email, title, employer
add column firmographic_fetched_at timestamptz; -- slow: headcount, stack, location
-- Rows whose volatile fields are older than the window, cheapest first.
select id, public_identifier
from people
where contact_fetched_at < now() - interval '45 days'
order by contact_fetched_at asc
limit 500;
The window is a policy, not a constant. Outbound lists that send tomorrow want the volatile clock tight, a month or less, because a bounced send costs more than a lookup. A dormant archive can let the same fields run for two quarters, because nothing reads them. Set the interval per segment and store it next to the segment, not in code. Feed the resulting due list through the same paced loop a nightly backfill uses, so a refresh sweep never outruns your rate budget.
Events beat the calendar
The cadence is the floor, not the strategy. The refreshes worth the most are the ones a calendar never triggers, because they follow a signal:
- A send bounces. The email is wrong now, whatever the timestamp said. Re-enrich on the bounce, not on the schedule.
- A reply says someone has moved on. That is a title and an employer change waiting to be fetched.
- A deal reaches a stage that needs firmographics you trust. Refresh the company at the moment the data is about to be used, not on the day a cron happened to run.
A blended policy, a loose per-field floor plus event triggers, re-fetches less than a tight calendar and carries fresher data at the point of use. The calendar catches the silent drift; the events catch the expensive surprises.
Store what you get back, so the diff is free
Event-triggered refresh only pays off if you can tell what actually changed, which means keeping the last response to compare against. That is the case for storing the raw payload: a refresh that overwrites in place tells you the new title but never that it changed, and the change is the part your pipeline wanted to act on.
The rule
Do not pick a refresh interval. Pick one per field group, set the volatile floor by how soon the data will be read, and let bounces and replies pull refreshes forward. You will re-fetch less and carry fresher data than any single number can give you.