AI
AI Chat Tools for Hiring: Screening and Interview Questions
Vendor pages document resume-file support, plan limits, data policies, and accuracy warnings for ChatGPT, Claude, and Gemini, but do not establish hiring effectiveness.
Sources checked 2 Oct 2026
For AI chat tools for hiring, [OpenAI’s File Uploads FAQ], Anthropic’s file-upload page, and Google’s file-upload help document file inputs that can carry resumes and job materials, but not screening accuracy, fairness, or hiring outcomes. Chat Picker has not tested ChatGPT, Claude, or Gemini for hiring; this guide compares documented features, limits, data handling, and a check you can run.
What the vendors document
The vendor details below are as read on October 1, 2026.
Files and plans. As read on October 1, 2026, OpenAI’s File Uploads FAQ says uploads are available on Free and paid plans, each file has a 512 MB hard limit, and Free users have three uploads a day. Anthropic’s upload page supports PDF and DOCX, sets a 500 MB chat-file limit and 30 MB project-file limit, and says PDFs of 100 pages or fewer include visual analysis while 101-to-1,000-page PDFs are text-only. Google’s upload help allows up to 10 supported files in one prompt, subject to availability, with a 100 MB limit for non-video files; work or school accounts may need an administrator to enable Drive access.
The main paid consumer plans listed are ChatGPT Plus at $20 per month, billed monthly, Claude Pro at $20 per month or $17 per month on annual billing, with $200 up front, and Google AI Pro at $19.99 per month. These subscription prices do not document cost per completed screen.
For interview research, OpenAI’s deep-research page says the feature can use uploaded files and the public web, with usage varying by plan. Anthropic makes Research available on Pro, Max, Team, and Enterprise, requires web search, and says sessions can use allowances faster. Google requires sign-in for Deep Research, includes Google Search by default, and lets users add files or connected Gmail and Drive sources.
Data, policy, and accuracy. For training on chats, OpenAI’s pricing page lists an opt-out on Free, Go, Plus, and Pro; Anthropic’s pricing page lists one on Free, Pro, and Max, while Team is not trained on by default. The Google AI plans page states no such setting, so Chat Picker’s method marks the item Not verified.
Policy boundaries also matter. OpenAI’s usage policies prohibit unauthorized aggregation, monitoring, profiling, or distribution of private or sensitive information, and prohibit inference about workplace emotions except for medical or safety reasons. Anthropic’s Usage Policy requires qualified professional review before covered advice, recommendations, or subjective decisions directly affecting people are finalized; direct model output also requires AI disclosure at the beginning of each session. Google’s guidelines reject factually inaccurate output that could cause significant real-world harm and statements advocating discrimination based on legally protected characteristics. These are vendor rules, not a legal compliance finding.
Each vendor also publishes an accuracy-related caution. OpenAI says ChatGPT can be incorrect or misleading and can sound confident when wrong, so it urges verification of important information. Claude’s help page says not to rely on it as the only source of truth and to inspect original cited sources. Google’s related-sources guide says source links may sometimes include public websites, uploaded files, and connected Workspace documents or email.
What the documentation cannot tell you
Vendor documentation does not report hiring-specific precision or recall, false-exclusion rates, bias-audit results, throughput, cost per screen, or reviewer agreement. It does not show how consistently an assistant will score equivalent candidates, whether generated questions measure the intended competency, or whether a particular resume layout survives upload and extraction.
Source links can support verification; they do not prove every inference is sound. You need a defined review process and documented decisions.
How to check it yourself
-
Resume screening. Give each assistant the same redacted resume and job description. Use: “Review the attached resume against the attached job description. For every requirement, quote the resume text that supports or contradicts it, label unsupported requirements, and do not infer skills that are not stated.” Look for evidence-based screening and missed qualifications. Record false exclusions, unsupported labels, omissions, and disagreements with trained reviewers. Treat these as your observations, not a vendor benchmark.
-
Interview question design. Give the assistant the job description, interview stage, and competency list. Use: “Create behavioral and technical interview questions that assess only the attached competencies. For each question, state the competency, the evidence it seeks, and an acceptable follow-up.” Look for job relevance and usable scoring cues. Record every irrelevant or potentially biased question and the rewrite you made.
-
Compliance and bias audit. Give each system the same rubric and de-identified candidate summaries. Use: “Audit the attached screening rubric and candidate summaries. Flag each item that is not job-related, asks for protected or irrelevant information, or turns an inference into a fact. Do not make the hiring decision.” Look for recommendation changes unsupported by job-related evidence. Record each flag and how a qualified reviewer resolved it.
-
Cost and throughput. Give identical files under the selected plans and keep tool settings as similar as the interfaces allow. Use: “Use only the attached resume and job description. Show evidence for every recommendation and list anything you cannot determine.” Look in the answer and session record for evidence handling, upload warnings, tool use, and context-limit messages. Record plan price, prompts, uploads, tool calls, limits encountered, elapsed time, and staff verification time; then calculate cost per human-reviewed screen from your records.
-
Scorecard. Use the same rubric and scoring scale. Try: “Apply the attached scorecard. Give each criterion a rating, cite the exact resume evidence, and identify missing information before making a recommendation.” Look for consistent ratings and clear uncertainty. Record rating changes, missing evidence, and agreement between reviewers.
Which rows of the comparison matter
On the Claude vs ChatGPT vs Gemini matrix, prioritize these rows:
- Access and price:
Free plan,Low-cost tier,Main paid plan,Heavy-use plans, andTeam plan. - Workload:
Context window in the appandHow usage limits are described. - Data handling and interface:
Training on your chatsandAds.
Check the vendor link and read date beside every figure. A Not verified blank means unknown, not zero; Chat Picker’s method says unconfirmed figures are left blank.