Case study · AI & Product

Designing an Evidence-First AI Recruitment Partner

The same career. Different resumes. Different conclusions about the person.

Role: concept, research, service design and workflow architecture · Status: Guide prototype in preview · Independent usability testing planned.

Everyone wanted a better application. I started questioning the starting point.

My LinkedIn feed kept bringing me two conversations about AI and recruitment that seemed to be talking past each other.

Recruiters were describing “AI slop”: applications polished into sameness, cover letters that repeated the advertisement and a disappearing sense of the human being behind the document. Some were even welcoming typos as a sign that a person had written it.

Meanwhile, candidates were being offered five prompts for resume success, Claude or ChatGPT rewrites, ATS keyword optimisation and apparently precise match scores. The promise was often another way to make an application faster or more persuasive.

Both sides were responding to real pressure. Candidates faced repetitive work and little visibility into selection. Recruiters faced application volume and weaker signals about who could actually do the job. These were observations from my feed, not a representative research study, but the tension was worth investigating.

The question became: what happens before the writing starts?

My own job search made the problem harder to ignore

I had several legitimate versions of my resume. A communications version brought one part of my career forward. A government and stakeholder version foregrounded another. Strategy and transformation, CX, product and digital each gave a different reader a useful view of the same experience.

When I used those documents with AI during my own job search, the system could reach different conclusions about the same person. Relevant capabilities would disappear from its assessment, then reappear when a conversation uncovered experience that was not on the uploaded pages.

I had changed the emphasis of the document. I had not changed my career.

That distinction mattered. If the system treated an omission as a lack of capability, it could screen out a suitable opportunity. If it filled the omission with a plausible assumption, it could write an application I could not honestly stand behind.

A resume is a lossy compression of a career

A resume is useful because it leaves things out. It selects the experience relevant to a particular audience and fits it into a document someone can quickly read.

A resume is a lossy compression of a career.

The better tailored that document becomes, the less suitable it becomes as a complete career database. Absence from the resume is not evidence of absence from the candidate.

This reframed the design problem. Before improving an application, the system needed a fuller account of the person, and a rule for what it could do with the things it did not know.

Recruiters infer. The difference is what constrains the inference.

An experienced recruiter does not read a resume as a literal inventory of every capability. Seeing “Head of Marketing” may reasonably prompt questions about budgets, agencies, procurement, planning, people leadership or executive reporting.

That professional context is useful. It generates hypotheses worth investigating. It does not establish which responsibilities this particular person held, at what scale, or with what results.

Human judgement also operates within reputation, networks, client relationships, accountability, feedback and commercial consequence. Those constraints do not make recruiters infallible, but they matter. AI can produce a similarly convincing inference without inherently carrying them.

“Act like an experienced recruiter” therefore left too much authority undefined. I needed a boundary between recognising a possibility and representing it as fact.

Professional context can generate the question.Only evidence can generate the claim.

The system could ask whether I had managed a budget. It could not write budget ownership into my biography because someone with my title probably had.

The candidate's task is bigger than writing a resume

Most candidates want to find work that genuinely fits, advance their career and put their strongest real self forward. They also want to spend less time on the routine work of recruitment. They are not trying to game their way into unsuitable jobs.

I recognised that administrative burden in my own behaviour. I had opened application portals asking me to re-enter my resume, looked at the time and left without applying.

The job to be done became:

Help me find work that fits my capabilities, career goals and circumstances; represent my strongest real experience; and reduce the repetitive work of searching and applying.

Continuous support from an excellent personal recruiter could help with much of this. Few candidates can afford that kind of ongoing representation. AI offered an opportunity to support parts of the task, provided the candidate retained direction and authority.

How might we?

How might we use AI to reduce the cognitive and administrative burden of job searching while preserving candidate truth, individuality and decision-making authority?

This led to a candidate-side AI recruitment partner: one that knows the candidate, knows what they want, searches routinely, assesses, asks, verifies, represents and preserves them. It automates routine work while leaving career decisions with the candidate.

AI is a force multiplier. It is not a compass. Here, that means using AI's range to examine opportunities, retrieve evidence and expose questions, while the candidate defines worthwhile work, acceptable compromises and what truthfully represents them.

Principles that shape the experience

Understand the career before tailoring the document. Build a source of evidence broader than whichever resume is convenient today.

Use context to ask, evidence to claim. Keep unknown information visible until the candidate supplies evidence or confirms its absence.

Separate ability from desire. “Could do this job” does not establish “should pursue this job”. Preferences must come from the candidate.

Preserve the person. Bring relevant evidence forward without inflating seniority, polishing away specificity or replacing a sound resume with a generic template.

Automate work around judgement. Research, retrieval and administration can be assisted. Direction, new factual claims, personal declarations and final submission remain human checkpoints.

Design the journey around understanding, then action

These journeys are conceptual models informed by observation and my own use, not measured descriptions of every candidate.

Current journey

Job boards → filtering → job description → selected resume → AI rewrite → missing keyword or inference → generic application → repeat.

Designed journey

Know me → know what I want → find opportunities → assess → ask → verify → represent me → preserve what was sent → prepare → learn. Routine work is assisted; career decisions stay with the candidate.

The application now follows a fuller understanding of the candidate and a deliberate assessment of the opportunity.

The audit extracts documented experience, then branches into “evidence shadows”. A yes leads to context, scope, contribution, outcome and recency. A no closes the branch. Limited exposure stays at its actual level. Verified evidence is retrieved rather than asked for again.

My historical audit reached approximately 500 sequential questions. The transferable method is adaptation to the person, not a 500-question template.

Governance that constrains the output

StateMeaningRequired behaviour
VERIFIEDSupported by candidate-supplied evidence.Use only within the scope established.
UNKNOWNNot currently established.Ask. Never silently turn it into fact or a gap.
GAPCandidate confirms absence.Record the specific gap. Never claim it.

Plausible inference → UNKNOWN → targeted question

  • Candidate evidence → VERIFIED.
  • Confirmed absence → GAP.
  • Unclear or unanswered → UNKNOWN.

AI cannot promote its own inference to VERIFIED. Evidence retains its source and strength; candidate confirmation does not imply independent verification.

A gate before drafting checks material requirements against the bank. The candidate may choose to proceed with an unresolved requirement, but the unsupported claim must remain omitted. Their decision to apply does not change its truth status.

Test the workflow before funding another product

My founder experience at People Wizards informed this decision. That work included psychometric recruitment tools as well as communication products, and took me through concept development, customer research, UX, product language, MVP work and go-to-market planning. It gave me another setting in which to examine AI's interpretation of career material, and a practical understanding of what it takes to operate an AI product. Dating.wtf and Workplace.wtf are no longer live.

The recruitment partner emerged from the combination of recruitment observations, my own job-search experience and that founder perspective. People Wizards informed the judgement about what to build; it was not the sole origin of the idea.

A hosted recruitment application would add inference costs, authentication, sensitive career-data storage, privacy and security responsibilities, infrastructure, maintenance and commercial overhead.

The candidate already has an AI environment. The first version can help them organise its available capabilities into a useful workflow, with manual alternatives where tools are missing.

The MVP decision was deliberate: test “Does this workflow create value?” before “Should this become software?” A free guided resource lets me investigate transferability before funding a separate application layer.

Operational feedback changed the design

Reviewing the guide against my working system exposed a gap: the draft explained the evidence method, but underrepresented the work the system was already helping perform.

That operational workflow includes live vacancy research, open-status checks, formatted .docx packs using approved templates, visual document inspection and assistance with application portals. It also has to handle document limits, interrupted authentication and failed portal attempts, while stopping for new claims, personal attestations and final submission.

The more consequential weakness was continuity. Too much knowledge remained distributed across conversations, documents and remembered corrections. Wrong contact details, repeated questions and uncertainty about previous submissions showed why conversational memory alone was inadequate.

These are observations from my own operational use, not independent usability findings. They led to a specific revision: give the workflow durable records as well as instructions.

Operational detail

Keep the inputs and records distinct
InputContentsPurpose
Career Evidence BankCapability, role/project, context, scope, actual contribution, outcome, recency, status, strength and source.Establish what may truthfully be claimed.
Recruiter configurationDirection, target roles, pay, work design, preferences, exclusions and acceptable compromises.Establish what is worth pursuing.
Market opportunitiesSourced vacancies, requirements, conditions, checked dates and unresolved details.Establish what needs assessing.

Evidence + preferences + vacancy → assessment → candidate questions and decisions → verification → application.

Interviews and outcomes feed back into the process. New candidate evidence updates the bank; approved preference changes update the configuration. Generated applications remain outputs, never independent evidence.

This separates “could do this job” from “should pursue this job”.

Seven records support execution and continuity:

RecordWhat it prevents
Authoritative candidate profileOld resumes overriding current contact details or confirmed routine facts.
Structured Career Evidence BankUnsupported claims and repeatedly rediscovering established experience.
Dated recruiter configurationLosing nuanced salary trade-offs, contract/fractional preferences, travel limits or specific exclusions.
Application ledgerDuplicate applications and confusion between prepared, submitted and interviewing.
Submitted-artifact archivePreparing from the wrong resume or overwriting what an employer received.
Portal-answer and permissions registerConfusing a known answer or standing preference with permission to disclose, attest or submit.
Document formatting standardRecreating approved documents with inconsistent typography, spacing or page dimensions.

A workspace index identifies the current files and versions. Each session reads those records first and leaves dated changes and a handover. Historical submissions stay unchanged even when current facts or preferences are corrected.

The candidate profile governs current identity and contact fields; the evidence bank governs capability claims; the configuration governs preferences. The archive establishes what was actually sent. A conflict is surfaced for resolution rather than settled by whichever document was uploaded last.

Make execution and its limits explicit

The design includes a capability check for live search, file creation, rendering, browser control, scheduling and durable storage. A text-only environment can support assessment and drafting; it cannot truthfully claim to have rendered a document or completed a portal. Each unavailable capability has a named manual handoff.

Portal assistance uses approved files and current routine answers within the candidate's authorisation. New facts, personal declarations and final submission remain candidate checkpoints. The permissions register documents these boundaries without overriding the tool's own requirements.

Submission needs evidence too. If a portal crashes or no confirmation appears, the ledger records uncertainty and checks for a receipt or an existing application before retrying. A click is not proof of submission. The archive preserves the exact sent files and answers, linked to the confirmation or a clearly labelled candidate report.

Interview preparation then retrieves that specific submission. Outcome analysis records the actual stage reached and genuine feedback. Repeated patterns may suggest something worth investigating, but one rejection cannot establish its cause.

Preserve the person, too

An application can be factually close to the source and still erase individuality through generic language, lost specificity or altered seniority.

The workflow therefore preserves a sound existing resume template where practical, including its hierarchy, typography, spacing and writing rhythm. Genuine parsing problems call for a specific diagnosis and the minimum correction.

QA has two layers: machine/process checks for readability and submission requirements, and human-credibility checks for supported claims, voice, specificity and accurate seniority.

Where tools allow, document production creates real .docx files and renders every page for inspection against the approved standard. Portal file limits may require a new version and another check. The exact reviewed version must be the one selected for upload.

The purpose is to ensure AI hasn't erased the human whose career it is describing.

The companion resource

The public guide translates this architecture into ten practical stages, with expandable instructions, copyable prompts and reusable record templates. It covers research, document production, portal assistance and preservation of the submitted application. The candidate works in their own AI account; the website supplies the method and setup materials.

Validation: can someone use it without me?

My friend was fascinated watching the workflow in practice. That interest provides a starting point for usability testing, not evidence of product-market fit.

The research question is:

Can another job seeker independently configure and use the system without Rebecca operating it for them?

The proposed first test follows a participant from workspace setup through evidence gathering, discovery, assessment, document production and a reviewed portal draft where tools permit. Actual submission remains their decision. A walkthrough can test the handoff without sending an unwanted application. Allow pauses and distinguish independent completion from completion with help.

Observe:

  • Where instructions confuse the participant or they abandon the workflow.
  • Whether the adaptive audit uncovers real omitted experience.
  • Whether they understand and correctly apply VERIFIED, UNKNOWN and GAP.
  • Whether search results fit their criteria better than their usual approach.
  • Whether application claims are supported and the document remains recognisably theirs.
  • Whether a fresh session retrieves current contact details, preferences and approved formatting without repeating resolved questions.
  • Whether it detects a duplicate or uncertain application, respects declaration/submission stops, and preserves the exact sent version after an authorised submission.

Record interaction problems without collecting confidential career material. Early findings would concern usability and perceived usefulness, not improved hiring outcomes.

Measurement and iteration

Proposed signalWhat it helps investigate
Page visits and repeat visitsReach and return use.
Step progression and abandonmentWhere readers continue or stop.
Prompt copiesWhich instructions readers take away.
Self-reported completionWhether readers say they completed the guide.
Observation and qualitative feedbackWhat helped, confused or required assistance.
Participant-observed record and handoff checksCorrect retrieval, repeated questions, version mistakes and recovery from interrupted work.

Copying a prompt does not prove it was used. Guide completion does not establish evidence accuracy. These signals need interpretation alongside observation.

I would use existing site analytics where appropriate. Aggregate tracking would need to be configured and checked; locally saved progress alone is not analytics. No career content is needed for these measures.

The website cannot observe private document production, portal completion or archive accuracy through prompt-copy events. Assess those through participant walkthroughs and consented, non-sensitive observations, not automatic collection of career records.

Findings will be added after independent usability testing.

Future updates will show what was observed, what changed and what happened on retesting. No independent findings are claimed yet.

AI is a force multiplier. It is not a compass.

Its value here is range: more opportunities to examine, evidence to retrieve, questions to ask and less administration to repeat. Direction, truth and career decisions remain with the candidate.