Data-Driven Recruitment, Sourcing First: 9 Metrics, Clean Data and the Decisions They Should Change

Data-driven recruitment is the practice of making sourcing, screening and hiring decisions from recorded evidence instead of memory. In practice that means tagging where every candidate came from, dating every stage change, scoring interviews against a fixed rubric, and agreeing in advance which number is allowed to change which decision.

Most teams start measuring at the wrong end. In Gem's 2026 Recruiting Benchmarks Report, job boards and company marketing produced roughly 90% of applications but only about half of hires, while direct sourcing produced 11% of hires from 2.6% of applications. The cheapest gains usually sit at the top of the funnel, in who you find and who replies, long before an offer is on the table.

Full disclosure: I run The Cognitive, which sources candidates across ~900M public profiles and interviews them live with AI. It produces sourcing data and per-candidate interview reports. It does not do funnel analytics, and this guide says exactly where that line sits.

Key takeaways

What is data-driven recruitment?

Data-driven recruitment, also called data-driven recruiting, data-driven hiring or data-driven talent acquisition, is hiring where each recurring decision has a number behind it. Which channel gets next month's sourcing hours. Which message gets rewritten. Which interview criterion gets dropped because it predicts nothing. The label covers the same work under any of those names.

The test I use is simple. If a metric moved by half tomorrow, would anyone do something different? If nobody would, move it to an appendix and keep the dashboard for numbers that trigger action.

Recruitment data comes in 4 kinds, and each answers a different question:

"Big data for recruitment" mostly means the first kind bought at scale: talent mapping data, salary sets from salary benchmarking tools, profile databases. Data analysis for recruitment is the work of joining all 4 kinds by candidate, so you can see that the people from one search, contacted with one message, scored well and stayed.

How do you build a data-driven recruitment strategy?

A data-driven recruitment strategy is a short written plan that fixes your definitions, your sources of truth and your review rhythm before anyone builds a chart. Mine has 5 steps, in this order:

  1. Write down the decisions. List the 5 or 6 calls your team makes every month, such as channel mix, message choice and interview loop design.
  2. Pick one metric per decision. Use the table in the next section. A decision with 3 metrics gets argued; a decision with 1 gets made.
  3. Name the system of record for each number. Source of hire lives in the ATS. Reply rate lives in your outreach tool. Never let 2 tools both claim the same number.
  4. Clean the inputs for 4 weeks before you read anything. The section on clean data below is the checklist.
  5. Review on 2 clocks. Sourcing and outreach numbers weekly, because you can change a search on Tuesday. Quality of hire and cost per hire quarterly, because they need time and volume.

Which recruitment metrics matter, and how do you calculate them?

These 9 are the ones I would track first. The third column is the point: each metric is listed with the decision it should change, so nobody tracks it for its own sake. The formulas are plain arithmetic. For what "good" looks like at each stage, use the dated benchmarks in our funnel conversion rates guide rather than any number repeated without a source.

MetricFormulaThe decision it changesWhere the data lives
Source of hireHires from one source ÷ all hires, × 100Where next quarter's sourcing hours and budget goATS source field, set at first touch
Sourcing yieldPeople shortlisted ÷ people a search returned, × 100Whether to fix the brief before sending a single messageYour sourcing tool
Contact reach rateContact lookups that returned an email or phone ÷ lookups tried, × 100Which channel to use for which kind of profileYour sourcing or contact tool
Reply ratePeople who replied ÷ people contacted, × 100, split by channel and messageWhich message to rewrite and which profiles to stop contactingYour outreach tool
Time to first interviewDays from first contact (or application) to the first completed interviewWhere scheduling slack is losing peopleOutreach dates plus interview dates
Stage conversionPeople who reached the next stage ÷ people who entered this stage, × 100Which stage is too loose, too tight or inconsistentATS stage history
Quality of hireYour written proxy, for example hires rated at or above expectations at 6 months ÷ all hires in the cohortWhich sources and which interview criteria predicted good hiresManager reviews and HR records
Offer acceptanceOffers accepted ÷ offers extended, × 100Pay, speed and how the role was soldATS offer records
Cost per hireExternal plus internal recruiting costs in a period ÷ hires in that periodChannel mix and tool spendFinance records plus ATS hires

2 of these need a written definition before you trust them. Quality of hire has no standard formula, so pick one proxy and hold it still for a year. Cost per hire breaks the moment someone adds recruiter salaries in one quarter and leaves them out in the next. If you want the longer stage-by-stage list, our 14 recruitment KPIs post has it.

How do you collect clean recruitment data?

Clean recruiting data is a by-product of a few boring rules that every recruiter follows every day. No analytics tool fixes inputs that were never recorded.

Fix the source field first

Give the source field a short, controlled list of 8 to 12 values and remove free text. Set it at first touch, never at hire, or every hire gets credited to whatever happened last. Split "sourced" from "applied" and give "rediscovered" its own value. Gem's report found that 46% of sourced hires now come from rediscovered candidates already in a CRM or ATS, up from 26% in 2021, and you can only see that shift if the tag exists. Our candidate rediscovery guide covers how to search that pool.

Date every stage change

Moving a card is not enough. Each candidate needs an entry date and an exit date for every stage, or stage conversion and time to first interview are guesses. Name stages after decisions, such as "hiring manager interview passed", not after activities, such as "chatted". Our funnel guide walks through renaming stages without breaking your history.

Record a reason on every exit

Keep a reason list of about 6 codes for rejections and 6 for withdrawals. "Withdrew: pay" and "Withdrew: took another offer" point to different fixes. Make the field required at the moment a candidate is closed out.

Tag people who arrive from outside tools

Candidates found in a sourcing tool often enter the ATS by hand, and that is where source data gets lost. The Cognitive does not write people it sourced into your ATS automatically. When you add one, set the source to something like "Sourced: The Cognitive" and note the search name, so source of hire still adds up at the end of the quarter.

Run a 15-minute weekly check

Once a week, pull 3 lists: candidates with no source, candidates in a stage with no entry date, and likely duplicates. Fix them by hand, and give the job to one named person so it happens in busy weeks too.

What should a recruitment KPIs dashboard look like?

A recruitment KPIs dashboard should read top to bottom in the order candidates move, with sourcing first and outcomes last. Here is the layout I would build. It is illustrative and holds no real numbers; the labels show what each panel contains.

  1. Panel 1, Sourcing (weekly, per role): searches run, people returned, people shortlisted, sourcing yield. One line per open role.
  2. Panel 2, Reach and replies (weekly): people contacted, contact reach rate, replies by intent (interested, needs information, not interested, opted out), reply rate split by channel and by message version.
  3. Panel 3, Process (weekly): a stage conversion strip for each role, the median time to first interview, and a list of candidates who have sat in one stage longer than your own usual time.
  4. Panel 4, Evaluation (per role, as interviews finish): interviews scheduled, completed and missed, plus the spread of scores on each rubric criterion.
  5. Panel 5, Outcomes (quarterly): source of hire, offer acceptance, the quality of hire proxy and cost per hire.

2 rules for every tile: print the count behind each rate (n), and print the one-line definition under the number. A rate with no n hides whether it came from 4 people or 400. If you need a recruitment dashboard sample to fill in each week, the recruiting metrics guide has a one-screen weekly report, and our recruitment analytics software review covers the tools that build the charts.

How do you use sourcing data to fix the top of the funnel?

Sourcing data shows which searches, profiles and channels turn into conversations, and it moves faster than any other recruitment data. You can change a search today and read the result tomorrow, which you cannot do with quality of hire.

Read sourcing yield per search before you send anything

If a search returns 50 people and you would shortlist 3, fix the brief before you touch the outreach. Change the filters before a single message goes out, because a weak list makes every downstream number look worse. In The Cognitive's AI sourcing tool, the brief is read into editable filters you can check. In my own test, "Senior backend engineer with 5+ years of Go who has built payments systems, Remote EU" came back with a location filter reading Remote EU as Germany plus 7 other countries, which is exactly the kind of line you want to see before judging the results.

Compare who you shortlist with who replies

Put shortlisted, contacted, replied and interviewed side by side for one role, then group by a profile attribute: current title, company type, years in role or location. Patterns show up quickly. Senior people at large companies might shortlist well and reply rarely, while people 2 years into a scale-up reply often. Shift the next search toward the group that replies and keep the shortlist bar where it is. If your outreach tool can't split replies by message version, our recruiting outreach software comparison shows which ones can.

Weight channels by hires, not by volume

Gem's 2026 figures are the case for this: referrals converted at 11 times the rate of inbound applicants (the case for a real employee referral program), internal mobility at 32 times, and sourced candidates were nearly 8 times more likely to be hired than inbound applicants. Your own ratios will differ, which is why source of hire needs the controlled tag from the previous section.

Let shortlist decisions train the ranking

Every Shortlist and Hide click is data about what you actually want, often more honest than the job description. In The Cognitive, once a role has 8 or more decisions, including 3 or more shortlists, what you shortlist and pass on shapes later ranking. Everyone found stays with the role, and later searches skip people you have already seen, so yield on the next search counts new people only. When the pool is thin, the search relaxes its softest filters and says what it broadened. Note that in your log, because it changes what the yield means.

How do structured interviews turn evaluation into data?

Structured interviews turn evaluation into data because every candidate for a role is scored on the same criteria and the same scale, so the scores can be compared. That comparability is also why they predict well. In the 2022 re-analysis by Paul Sackett and colleagues Zhang, Berry and Lievens in the Journal of Applied Psychology, structured interviews had the highest mean operational validity of the selection procedures reviewed, at .42, as summarised by SIOP's TIP.

Unstructured interview notes can't be compared across candidates, so they can't be analysed. If you want evaluation data, record 3 things for each candidate: a score per criterion on a fixed scale, the evidence behind it in writing, and the verdict kept separate from the score. After 2 or 3 quarters you can check which criteria actually line up with your quality of hire proxy, and drop the ones that don't.

The Cognitive's AI interviewer runs this as a live, two-way AI video interview of 10 or 20 minutes, in 9 languages. Candidates book their own slot without an account, and the invitation tells them the interview is run by AI. The rubric is fixed per role and the questions adapt live to each answer. Each report gives a 1 to 5 score per criterion, a weighted score out of 100, a suggested verdict, overall written feedback, the transcript and the recording. It also checks up to 5 resume claims and marks each one verified, refuted or unclear, with evidence from the interview. Nothing is rejected automatically, and integrity flags are logged, not scored.

This 5-minute recording shows a live AI interview for a go-to-market role. The AI asks the candidate about an outbound email campaign they ran, then follows up on reply rates and on how they qualified leads, adjusting each question to the previous answer.

What are the pitfalls of data-driven recruiting?

Data-driven recruiting goes wrong in 5 predictable ways. Each one has a cheap fix.

Vanity metrics

Applications received, messages sent and profiles viewed all rise with effort and say little about hiring. Gem's split of roughly 90% of applications against about half of hires shows how far volume and outcome can drift apart. Replace each volume number with the rate that follows it: applications become stage conversion, messages sent become reply rate.

Small samples

An offer acceptance rate of 50% from 2 offers is 1 decline. My rule: never show a rate without its n, and never act on a rate built from fewer than about 20 events. Pool similar roles or wait a quarter instead.

Bias in historical data

Past hires reflect past decisions, including the biased ones. A model, a scoring rule or a recruiter's instinct trained on who got hired before will repeat those patterns. Check pass rates by group at every stage, not only at the end. Ranking that learns from your shortlists will learn your habits too, which is one more reason to audit those shortlists. Our guide to AI bias in hiring covers how to run that check, and The Cognitive has not published a bias audit.

Optimising one number

Push time to first interview alone and the bar quietly drops. Pair every speed metric with a quality metric and review them on the same page, so a gain on one that costs the other is visible the week it happens.

Someone else's benchmark

Benchmarks come from other teams' stage names and other roles. They are useful context and a poor target. Our funnel conversion rates guide dates and sources every figure it shows, and your own trailing 12 months, split by role family, is the comparison that fits.

Where does The Cognitive produce recruitment data?

The Cognitive produces sourcing data and per-candidate interview reports. It does not replace your ATS reporting or a recruitment analytics tool. Here is what it records, so you know what you can pull from it:

What it does not do matters just as much. It is not an ATS: no job posting, no offers and no funnel analytics. It does not search your ATS history, so the rediscovered candidates Gem counts still live in your own system. It connects to 60+ ATS platforms, and for candidates imported from your ATS it writes back a note with the score, a summary and the report link. People sourced inside The Cognitive are not written into the ATS automatically, which is why the tagging rule above matters. AI Sourcing starts at $49/month and AI Interview at $99/month, on monthly plans you can cancel anytime.

New accounts get a free trial with 100 sourcing credits.

Who is The Cognitive built for?

It is built for recruiters and founders who want better data at the top of the funnel: a search they can check, a shortlist that teaches the ranking, and interview scores they can compare across every candidate for a role. It suits teams that fill roles by sourcing rather than waiting for applicants, and teams whose interviewers are the bottleneck.

Measure the top of your funnel on one role Write the brief in plain English, see sourcing yield per search, and interview the people you shortlist live with AI. Start free

How should you start with data-driven recruitment this month?

Start with one open role for 4 weeks, not the whole team. Week 1, lock the source list and the stage names, and write the 9 definitions on one page. Week 2, run your searches and log yield per search before any outreach. Week 3, send outreach and track reply rate by message and by channel. Week 4, read the numbers with their n beside them and make one change: a new brief, a new message or a new channel. Then do it again on the next role.

If you want sourcing and interview data from the same place while you do it, sign up for The Cognitive and run that first role there.

Run one role the data-driven way Search ~900M public profiles from a plain-English brief, reveal contacts only when you need them, and get a scored report for every interview. Start free

Sources

Frequently Asked Questions

What is data-driven recruitment?

Data-driven recruitment is hiring where each recurring decision, such as channel mix, outreach message or interview design, is made from recorded evidence instead of memory. It rests on 4 kinds of recruitment data: sourcing, process, evaluation and outcome. Data-driven recruiting, data-driven hiring and data-driven talent acquisition all describe the same practice. The Cognitive covers the sourcing and evaluation parts: sourcing yield per search and a scored report for every AI interview.

How do you use data in recruitment?

Use data in recruitment by naming the decisions first, then picking one metric per decision and one system of record per metric. Clean the inputs for 4 weeks, then review sourcing and outreach weekly and quality and cost quarterly. Start with one open role rather than the whole team, and make one change per cycle so you can tell what moved the number.

What is a data-driven recruitment strategy?

A data-driven recruitment strategy is a short written plan that fixes definitions, sources of truth and a review rhythm before anyone builds a chart. It lists the decisions the team makes every month, assigns one metric to each, names where each number lives, and sets 2 review clocks: weekly for sourcing and outreach, quarterly for quality of hire and cost per hire.

What should a recruitment KPIs dashboard include?

A recruitment KPIs dashboard should run in the order candidates move: sourcing, reach and replies, process, evaluation, then outcomes. Sourcing shows people returned, shortlisted and sourcing yield per role; outcomes show source of hire, offer acceptance, quality of hire and cost per hire each quarter. Print the count behind every rate and a one-line definition under every number.

What recruitment data should I collect first?

Collect 3 things first: a source tag from a controlled list set at first touch, an entry and exit date for every stage, and a reason code on every rejection or withdrawal. Those 3 fields make source of hire, stage conversion and time to first interview computable. Give rediscovered candidates their own source value; Gem found they made up 46% of sourced hires in its 2026 report.

How do you measure sourcing yield and source of hire?

Sourcing yield is people shortlisted divided by people a search returned, times 100; source of hire is hires from one source divided by all hires, times 100. Yield is read per search and changes this week's brief. Source of hire is read quarterly and changes where sourcing hours go. In The Cognitive, every search reports the number of candidates found and records each Shortlist and Hide, so yield per search is countable.

What is a good reply rate for recruiting outreach?

A good reply rate is one that beats your own trailing rate for the same kind of role, channel and message, because published reply-rate figures depend heavily on who was contacted and how. Track it split by channel and by message version, and show the count behind it. For dated stage benchmarks further down the funnel, see our recruitment funnel conversion rates guide.

Are structured interviews really more data-driven than unstructured ones?

Yes. Structured interviews score every candidate for a role on the same criteria and scale, so the results can be compared and analysed. In Sackett, Zhang, Berry and Lievens' 2022 re-analysis in the Journal of Applied Psychology, structured interviews had the highest mean operational validity of the procedures reviewed, at .42. The Cognitive's live AI interview keeps the rubric fixed per role while the questions adapt live, and scores each criterion from 1 to 5.

Does The Cognitive replace my ATS reporting or recruiting analytics?

No. The Cognitive produces sourcing data and per-candidate interview reports, and it does not do funnel analytics, job posting or offers. Keep your ATS as the system of record and your ATS reports or analytics tool for funnel charts. The Cognitive connects to 60+ ATS platforms and writes a note with score, summary and report link back for candidates imported from the ATS; people sourced inside it are not written into the ATS automatically.

How many hires do you need before recruiting data means anything?

You need enough events behind each rate, and my rule is about 20 before acting on any one number. An offer acceptance rate from 2 offers moves by 50 points on a single decline. Pool similar roles, wait a quarter, and always show the count next to the rate. Sourcing and reply data reach that volume fastest, which is another reason to start data-driven recruitment at the top of the funnel.

Related reading

All posts · thecognitive.io