TryCrucible candidates complete real hands-on challenges scored across 6 dimensions. Browse verified profiles, filter by skill and score, and reach out directly. No resume screening required.
Verified skills, not claimed ones
Every candidate on TryCrucible has completed at least one hands-on challenge evaluated by our AI and optionally a human expert. The score is attached to real code you can inspect.
Cut screening time drastically
Filter by challenge category and minimum score before you ever look at a resume. You reach out only to candidates who have already proved they can do the job.
See the work, not the CV
Each profile links directly to the submission repo. Read the decisions doc, browse the code, check the AI score breakdown — all before a first call.
Targeted to AI engineering
RAG pipelines, agents, MCP servers, evals, coding agents, AI tool proficiency. Every challenge maps directly to the skills modern AI teams actually need.
A common question from hiring teams: “what stops a candidate from having someone else do the work?” Several overlapping signals make proxy work detectable — and hard to hide.
Temporal minimums
Every challenge has a minimum completion time based on difficulty. Submissions that arrive suspiciously fast are automatically flagged — there is no way to "paste in" a finished solution and pass.
GitHub provenance checks
The submission repo is cloned and inspected before evaluation begins. Repo age, commit count, author email, and commit timestamps are all checked. A repo created 30 minutes ago with one commit does not look like a week of real work.
Decisions doc authenticity
Every submission requires a decisions.md explaining key engineering choices. Our AI evaluates whether the explanations are specific to the candidate's own implementation — vague, generic, or corporate-sounding prose is flagged for human review.
Cross-submission similarity
Submissions are embedded and compared using vector similarity. Near-identical vectors across different candidates — same challenge, same decisions text, same code structure — are caught automatically.
Human review for high scores
Any submission scoring 88 or above is routed to a human expert reviewer from our network before the score is finalised. The reviewer reads the code, the decisions doc, and the AI evaluation. A proxy's work rarely survives this.
Scoped LLM keys, not internet access
Candidates work with a temporary LLM key scoped only to their challenge. The sandbox has no internet access — no Stack Overflow, no ChatGPT, no external APIs. They can only use the provided model and the challenge dataset.
No system is perfect — but the bar is high
No hiring signal is cheat-proof. But outsourcing a challenge to a proxy means the proxy also needs to: build the project, write a plausible decisions doc specific to that code, submit within realistic time windows, and pass human expert review for high scores. That combination catches the vast majority of bad-faith attempts — and makes the rare ones expensive enough to deter.
Beyond candidate search — a full set of tools built specifically for AI engineering hiring teams.
Private company challenges
Create custom challenges targeting your exact tech stack. Only candidates you invite can see or attempt them. Set time limits and custom rubrics.
Direct & bulk invitations
Invite candidates one-by-one, upload a CSV of up to 200 emails, or share a single open-link. Track acceptance and completion status in real time.
Talent pipeline
A built-in CRM. Move candidates through Interested → Contacted → In Process → Hired stages alongside their verified scores — all in one view.
Score-threshold alerts
Set a category and minimum score. Get notified the moment any candidate clears your bar — no manual checking, no missed talent.
ATS webhook
Register a webhook URL and receive a signed payload instantly when an invited candidate is scored. Plug into your existing ATS or Slack workflow.
Analytics dashboard
Full invite funnel: invited → accepted → submitted → scored. Average scores, time to completion, and drop-off rates per challenge.
Scores come from a combination of automated test execution and GPT-4o evaluation. Every submission above 88 also gets a human expert review before the score is published.
Correctness
The submission is run against real test inputs. Output is compared to ground truth. You can inspect the exact inputs and the candidate's outputs — nothing is hand-waived.
Architecture
AI reads the code and evaluates structure, separation of concerns, and whether the approach is appropriate for the problem. Spaghetti code scores low regardless of whether it passes tests.
Decision Quality
Every candidate must write a decisions.md explaining key engineering choices. The AI checks whether the explanations are specific to their implementation — generic or ChatGPT-sounding prose is flagged.
LLM Usage
Candidates work with a provided LLM key. AI evaluates whether they used it effectively — good prompt design, appropriate tool usage, avoiding hallucination errors. This is a core AI engineering skill.
Robustness
Edge cases, error handling, graceful degradation. The runner throws unexpected inputs at the system. A brittle solution that only works on happy-path data scores poorly here.
Clarity
Is the code and documentation readable? A future hire must maintain this code. The AI evaluates naming, comments where needed, and README quality.
Score range: 0 – 100
Industry average across all submissions is ~68. A candidate scoring 85+ is genuinely strong. We recommend filtering by 75+ as a starting point for senior roles.
Create a company account. Candidate search is included from day one.
Filter by challenge category and minimum score. Browse verified profiles instantly.
Check submission repos, score breakdowns, and AI evaluation notes before reaching out.
Contact candidates directly. No middlemen, no recruiter fees.
Every candidate has a verified score. Every score is backed by real, runnable code. Free trial — 10 searches, no card required.
Questions? Email us at raj@webpipl.com