← thomascha.com

Case Study 02

AI-Powered Job Fit Evaluation System

Applying a Research-First Methodology to an Unfamiliar Domain

Workflow AutomationOccupational PsychologyPrompt Architecture

Live output

This site is a live output of this system. Every role that led to thomascha.com was screened through the evaluation framework described here. That includes the postings that scored too low to pursue, which shaped what you're reading now.

See a full evaluation from a real posting →

The Problem

Most job fit tools measure skills and personality type. Person-environment fit research identifies a more predictive variable set: authority structure, feedback predictability, urgency culture, collaboration intensity, and the specific type of ambiguity a role requires. Standard tools don't screen for those. Manual evaluation doesn't either — it's slow, inconsistent, and produces no documentation a second reader could use.

For a career transition in an emerging field, the problem compounds. Roles appear under inconsistent titles across platforms. Postings are written in HR language that obscures what the job actually involves day to day. And there's no reliable way to evaluate fit without investing significant time in an application first. That cost lands twice: once on the applicant, once on the employer.

The System

The same three-stage methodology from Case Study 1 — research thread, prompt architecture, deployment — applied in a domain with no prior expertise.

Phase 01

Requirements and Fit Assessment

A role snapshot strips HR language and describes what the job actually involves day to day. A fit score with a one-line rationale follows. An alignment section maps specific posting requirements to documented strengths with explicit connections. A divergence section separates hard mismatches (genuine friction likely to filter you out) from soft gaps (closable in practice or overstated in the posting). Phase 1 closes with a candidate tier assessment against a three-level readiness framework.

Phase 02

Psychological and Environment Fit

A day-to-day fit check evaluates whether the described responsibilities would suit or drain you repeated five days a week. An 8-signal environmental red flag screen checks posting language for patterns that predict poor fit but rarely appear in skills-based evaluation. Absent signals are noted explicitly — their absence is also information. A vagueness assessment flags postings that can't describe what the role does, treating that inability as a risk signal in its own right.

Phase 03

Cross-Posting Summary

A compact six-field summary: role, fit score, key strengths for this posting, key gaps or red flags, verdict, and application status. Designed to be collected across multiple evaluations so postings can be compared side by side. When fifteen postings have been evaluated, you're comparing structured data rather than trying to remember what you thought three weeks ago.

What the System Catches

Phase 2 screens for eight named signal types in posting language. These are the variables person-environment fit research identifies as predictive and that standard tools consistently miss. Four are shown here.

Urgency Culture vs Deadline-Driven Work

Chronic-crisis operating rhythm vs. genuine project deadlines — the fit implications are opposite.

The system looks for phrases like "startup pace," "no two days are the same," "move fast," and "constantly shifting priorities." These indicate a chronic-urgency environment, where the operating rhythm is defined by reactive response rather than planned execution. The evaluation distinguishes this from genuine deadline-driven work (defined end dates, clear deliverables, finite sprint cycles), which is categorically different in its fit implications. A posting that describes both simultaneously triggers a deeper read of which mode is primary.

Feedback Predictability

Whether performance is evaluated against defined criteria or a shifting standard.

Postings that offer only vague performance markers ("cultural fit," "strategic impact," "stakeholder satisfaction") with no concrete output definitions signal environments where feedback can be arbitrary or withheld. This correlates with higher neuroticism load and is one of the strongest predictors of burnout in the research literature, stronger than workload alone. The evaluation checks whether the posting names specific success criteria and whether the review process is described in structural terms.

Performance Ambiguity

Roles where deliverables have no finish line and success criteria are left undefined.

Posting language like "help shape our direction," "define what success looks like," or "ambiguous and fast-moving environment" can signal roles where the scope, output, and evaluation criteria are all undefined. The evaluation distinguishes performance ambiguity from exploratory ambiguity (open-ended research, building something new with loose constraints) because the two look similar on the surface but have opposite fit implications. A role with genuine creative latitude is evaluated differently from one where vague expectations are a structural feature. This distinction is described in the Key Architectural Decision section below.

Decision-Making Authority Framing

Whether the role positions the candidate as a final authority or independent operator.

Phrases like "own the vision," "you'll build this from scratch," and "define the strategy" often frame the candidate as the primary decision-maker and public face of a function. For profiles that perform at their best with structured direction and execution authority (rather than as originators), this framing is a fit risk even when the work itself aligns well. The system flags it explicitly rather than reading it as an enthusiasm signal.

All eight signals are applied to a real posting in the sample evaluation.

The Iteration Arc

Getting the evaluation prompt right took more than editing. Three models — Fable, Opus, and Sonnet — ran the same evaluation against the same data independently. Each model then critiqued its own output and the other two. That cross-critique surfaced four failure modes common across all three: invention (fabricating supporting evidence), flattening (treating all profile dimensions as equally weighted), underselling (not naming genuine strengths directly), and palatability softening (hedging honest negatives to sound encouraging). The strongest elements from all three evaluations were synthesised into a single benchmark.

He set the target explicitly: match Fable's evaluation quality while running at Sonnet's cost. Fable produced the reference evaluation. Sonnet ran five rounds against it, with each round's shortcomings converted into specific, mechanical prompt rules.

Step 1 — Establish the benchmark

Fable
Opus
Sonnet
cross-critique produces 4 failure modes
Benchmark evaluation

Step 2 — Iterate to match

80%First Sonnet run against the Fable benchmark
85%Flattening corrected — dimensions weighted by profile priority
~90%Invention reduced — evidence must trace to posted facts
93%Underselling and palatability softening addressed
95%Prompt compliance confirmed

The remaining gap is model judgment — calibration on secondary dimension scores and mid-list hierarchy ordering — not prompt compliance. The system runs at 95% of benchmark quality at Sonnet cost.

Key Architectural Decision

Standard framing treats ambiguity tolerance as a single spectrum. The evaluation needed to distinguish two separate types. Performance ambiguity covers unclear expectations, undefined success criteria, and deliverables with no finish line. Exploratory ambiguity covers open-ended research, building something new with loose constraints, and tinkering on an undefined problem. A posting that signals chronic performance ambiguity and one that signals genuine creative latitude can look identical on the surface. The fit implications are opposite, and collapsing them into a single variable would have made the evaluation actively misleading.

This distinction came from the research phase, not from editing the prompt. A prompt built from memory would have reproduced the standard framing: ambiguity tolerance as a single dial. Mapping the occupational psychology literature before writing a single evaluation criterion surfaced it. That's the same reason the research thread in Case Study 1 came before the prompt. The constraints that make a system accurate aren't always the ones you'd think to include on any given day. Having AI research them first means the output is more thorough than recall alone.

The So What

The problem this system solves isn't specific to job hunting. Any decision process that requires evaluating an external opportunity against a complex personal or organizational profile — and producing consistent, documented output each time — faces the same bottleneck. The manual version is slow, inconsistent, and dependent on the evaluator's recall and state of mind. A research-grounded AI system produces the same quality of analysis every time, documents its reasoning, and scales without degradation.

The system has grown past its original scope. A 17-dimension profile framework drawn from I/O psychology and burnout research — covering dimensions that commercial tools consistently miss, including values congruence, feedback delay tolerance, and interest congruence — now forms the foundation of a product design in active development. A 23-question intake has been built and live-tested across three models with the same respondent. The personal version proved the concept. The product is the next step.

He researched the I/O psychology and burnout research literature from scratch, encoded it into a three-phase evaluation system, iterated the prompt through multi-model comparison and five rounds of benchmark-driven refinement, built a 17-dimension intake framework, tested a 23-question intake across three models with live subject runs, and deployed the system against 40+ real job postings. No computer science background. No prior expertise in hiring psychology.

Sample Output

See the system's full output on a real posting.

A complete Phase 1 and Phase 2 evaluation — role snapshot, fit score, alignment and divergence analysis, 8-signal red flag screen, and honest verdict — run on a real job posting at 7.5/10 fit.

Read the sample evaluation →