Skip to main content
For hiring teams

Hire AI engineers who have already proved it

TryCrucible candidates complete real hands-on challenges scored across 6 dimensions. Browse verified profiles, filter by skill and score, and reach out directly. No resume screening required.

// Why TryCrucible

A better signal than any resume

Verified skills, not claimed ones

Every candidate on TryCrucible has completed at least one hands-on challenge evaluated by our AI and optionally a human expert. The score is attached to real code you can inspect.

Cut screening time drastically

Filter by challenge category and minimum score before you ever look at a resume. You reach out only to candidates who have already proved they can do the job.

🔍

See the work, not the CV

Each profile links directly to the submission repo. Read the decisions doc, browse the code, check the AI score breakdown — all before a first call.

🎯

Targeted to AI engineering

RAG pipelines, agents, MCP servers, evals, coding agents, AI tool proficiency. Every challenge maps directly to the skills modern AI teams actually need.

// Integrity & verification

How we keep scores honest

A common question from hiring teams: “what stops a candidate from having someone else do the work?” Several overlapping signals make proxy work detectable — and hard to hide.

Temporal minimums

Every challenge has a minimum completion time based on difficulty. Submissions that arrive suspiciously fast are automatically flagged — there is no way to "paste in" a finished solution and pass.

🔗

GitHub provenance checks

The submission repo is cloned and inspected before evaluation begins. Repo age, commit count, author email, and commit timestamps are all checked. A repo created 30 minutes ago with one commit does not look like a week of real work.

📝

Decisions doc authenticity

Every submission requires a decisions.md explaining key engineering choices. Our AI evaluates whether the explanations are specific to the candidate's own implementation — vague, generic, or corporate-sounding prose is flagged for human review.

🧬

Cross-submission similarity

Submissions are embedded and compared using vector similarity. Near-identical vectors across different candidates — same challenge, same decisions text, same code structure — are caught automatically.

👤

Human review for high scores

Any submission scoring 88 or above is routed to a human expert reviewer from our network before the score is finalised. The reviewer reads the code, the decisions doc, and the AI evaluation. A proxy's work rarely survives this.

🔐

Scoped LLM keys, not internet access

Candidates work with a temporary LLM key scoped only to their challenge. The sandbox has no internet access — no Stack Overflow, no ChatGPT, no external APIs. They can only use the provided model and the challenge dataset.

No system is perfect — but the bar is high

No hiring signal is cheat-proof. But outsourcing a challenge to a proxy means the proxy also needs to: build the project, write a plausible decisions doc specific to that code, submit within realistic time windows, and pass human expert review for high scores. That combination catches the vast majority of bad-faith attempts — and makes the rare ones expensive enough to deter.

// Company toolkit

Everything you need to run your hiring

Beyond candidate search — a full set of tools built specifically for AI engineering hiring teams.

🏗️

Private company challenges

Create custom challenges targeting your exact tech stack. Only candidates you invite can see or attempt them. Set time limits and custom rubrics.

📧

Direct & bulk invitations

Invite candidates one-by-one, upload a CSV of up to 200 emails, or share a single open-link. Track acceptance and completion status in real time.

📊

Talent pipeline

A built-in CRM. Move candidates through Interested → Contacted → In Process → Hired stages alongside their verified scores — all in one view.

🔔

Score-threshold alerts

Set a category and minimum score. Get notified the moment any candidate clears your bar — no manual checking, no missed talent.

🔗

ATS webhook

Register a webhook URL and receive a signed payload instantly when an invited candidate is scored. Plug into your existing ATS or Slack workflow.

📈

Analytics dashboard

Full invite funnel: invited → accepted → submitted → scored. Average scores, time to completion, and drop-off rates per challenge.

// How scoring works

Six dimensions. One honest score.

Scores come from a combination of automated test execution and GPT-4o evaluation. Every submission above 88 also gets a human expert review before the score is published.

⚙️
25%

Correctness

The submission is run against real test inputs. Output is compared to ground truth. You can inspect the exact inputs and the candidate's outputs — nothing is hand-waived.

🏗️
20%

Architecture

AI reads the code and evaluates structure, separation of concerns, and whether the approach is appropriate for the problem. Spaghetti code scores low regardless of whether it passes tests.

🧠
20%

Decision Quality

Every candidate must write a decisions.md explaining key engineering choices. The AI checks whether the explanations are specific to their implementation — generic or ChatGPT-sounding prose is flagged.

🤖
20%

LLM Usage

Candidates work with a provided LLM key. AI evaluates whether they used it effectively — good prompt design, appropriate tool usage, avoiding hallucination errors. This is a core AI engineering skill.

🔒
10%

Robustness

Edge cases, error handling, graceful degradation. The runner throws unexpected inputs at the system. A brittle solution that only works on happy-path data scores poorly here.

📖
5%

Clarity

Is the code and documentation readable? A future hire must maintain this code. The AI evaluates naming, comments where needed, and README quality.

Score range: 0 – 100

Industry average across all submissions is ~68. A candidate scoring 85+ is genuinely strong. We recommend filtering by 75+ as a starting point for senior roles.

// Skills you can filter by

Every domain of modern AI engineering

⚙️ RAG Pipelines🤖 AI Agents🔌 MCP Servers💻 Coding Agents🧪 Evals & Testing🛠️ AI Tool Proficiency
Browse verified candidates →
// How it works

From sign-up to first contact

01

Sign up

Create a company account. Candidate search is included from day one.

02

Search & filter

Filter by challenge category and minimum score. Browse verified profiles instantly.

03

Review work

Check submission repos, score breakdowns, and AI evaluation notes before reaching out.

04

Reach out

Contact candidates directly. No middlemen, no recruiter fees.

// FAQ

Common questions

// Ready to find your next hire?

Stop reading resumes.
Start reading code.

Every candidate has a verified score. Every score is backed by real, runnable code. Free trial — 10 searches, no card required.

Questions? Email us at raj@webpipl.com