
After Resumes Are Templated, What Do Recruiters Actually Look For?
24 templates, paste the job description once, one chat window, from resume and application to interview. When this Apply AI Show HN came out, my first reaction wasn't about efficiency, but rather how it breaks down the job search into several measurable steps: input JD, generate draft, switch templates, enter interview prep. Is this direction good for publishing papers? Honestly, looking at the product alone, it's hard to publish, but the underlying representation learning is worth discussing. A resume is originally a low-dimensional projection of a person's experience, now re-encoded by the model again, like taking a PDF, doing OCR, then summarizing, then style transfer based on the target role.
The application page in the screenshot looks like a dashboard, but it's actually more like the frontend of a recruitment funnel. Left side is jobs, middle is chat, right side is resume status. If we put it back in an academic context, it's close to a query-document pair: JD is the query, resume is the document, and the model is responsible for rewriting the document to be closer to the query. Traditional ATS approaches lean towards lexical matching—keywords, job levels, years of experience, schools—to filter out mismatches. After large models are added, matching becomes narrative matching; the model adds verbs, adds achievements, adds "participated," "drove," "implemented." Vaswani et al.'s 2017 paper on attention mechanisms is interesting here: the model no longer just looks at local tokens but treats the entire JD as context to decide which experiences should be highlighted. The problem lies here too: the more complete the context, the smoother the generated content, making it easier to blur boundaries.
I prefer comparing applicant-side tools like Apply AI with recruiter-side tools like Zee. The former assumes candidates lack expression skills and need AI to organize their experiences into job language; the latter assumes recruiters lack time, so AI conducts an initial 20-30 minute interview first, then provides a shortlist to the team. These two paths seem complementary but actually point to the same metric: interview rate. As long as the metric is interview rate, the system will optimize for "making people want to click." Goodhart's Law fits perfectly here: when a measure becomes a target, it ceases to be a good measure. Candidates learn which words pass, recruiters learn which questions filter, and both sides are pushed by the same loss function.
I work in CV, and recently processed a batch of PDF resumes. It took a while before I realized the bottleneck isn't whether the model can write, but whether the evidence chain can be verified. A bullet point saying "improved recognition accuracy by 12%" would make reviewers ask about baselines, datasets, test set leakage, and confidence intervals. Recruiters won't ask that deeply, but the essence of the problem is similar. LLMs can generate "better" verbs, but they can't prove you did it. Bender et al.'s 2021 critique of stochastic parrots holds true for job-seeking AI too: the model samples from text distributions, generating fluent, human-like, job-like text, but it doesn't inherently know facts. The more templated the resume, the more HR sees what looks like multiple submission versions of the same paper.
If I were designing an experiment, I wouldn't just look at "whether interview rates increased after using AI." That's too easy to self-validate. The control group needs at least three types: human-written resumes, templated resumes, LLM-generated resumes; job types must be stratified (engineering, product, academic, operations) because feedback cycles differ; also need to control for application time, job popularity, and candidate background. Real error analysis needs to be broken down: was it keyword hits, exaggerated experiences, or did recruiters just like a certain tone? I wrote a post last week about municipal AI dashboards, saying the biggest lack is public error analysis; it's similar here. If recruitment AI only reports pass rates without reporting false positives and false negatives, it's optimizing for clicks, not matching.
When lab funding is tight, I often think of these tools as spot GPUs: cheap, runnable, suitable for exploration, not for conclusions. Apply AI has value for ordinary job seekers; it lowers the barrier to expression, allowing those who aren't good at writing to still organize their experiences.
But its risks are obvious: resumes shift from "personal archives" to "job response functions." Zee faces similar risks; AI initial interviews can expand screening but turn interviews into standardized Q&A, where silence, nervousness, and atypical expressions might be misjudged. Whether a role really needs someone who can chat or someone who gets things done, the model may not distinguish clearly.
So I care more about one experimental metric: what exactly are both parties exchanging—signal or noise? If AI makes resumes prettier and interviews more performative, does the recruitment system become more accurate, or just more expensive?
📌 This article is compiled from Hacker News, original source: https://apply-ai.work
Copyright belongs to the original author. This article is a compilation and independent analysis based on public reports.
Physix Frontier