Read authentic reviews from customers, clients and employees.
Find top rated recruiters on the GR Marketplace
 

Benchmark IT - Technology Talent

4.78 from 1262 reviews
 
Job
Senior Technical Lead, Autonomous-Agent Verification, Long-Term Contract
Upper Saddle River, New Jersey, United States
CONTRACT
Our direct client, a global professional-services firm, is building autonomous AI agents for use in regulated financial work. They need a hands-on lead to design the independent verification layer. It sits outside the agent teams and decides whether an agent's behavior can be relied on. The client's position is that trust is the product: every result must be traceable to its evidence, and every limit must be documented. You build the system that tries to break, measure, and certify the agents. You don't build the agents themselves. Note: this work can be performed from home, anywhere in the US (most meetings will be on East Coast time).   It is also open as a full-time salaried position as an alternative to long-term contract.   What you'll do
  • Design the "referee" system. Build repeatable, isolated test environments in which agents act against simulated tools that keep their own state. Wire the results into release pipelines so a failing score blocks a ship decision. The referee's own outputs (scores, evidence, verdicts) must be versioned and able to withstand outside scrutiny.
  • Decide what "correct" means for multi-step work. Static benchmarks aren't enough. You'll define how to judge:
    • whether numbers and facts match known-true data, including a tolerance for what counts as a meaningful error rather than a trivial one
    • the whole journey of a run: how it planned, which tools it chose, how it recovered from mistakes, and whether it wandered from its goal
    • whether each conclusion can be followed back to the specific source that supports it
    • whether behavior stays within accounting and regulatory rules
    • whether the measurements themselves are statistically sound, including how many runs are enough, how to separate real regressions from noise, and the difference between "succeeds at least once in k tries" and "succeeds every time in k tries"
  • Lead the team. Set the technical direction, mentor engineers, and set the standard for the test infrastructure.
  • Build and calibrate model-based graders. When one AI grades another, you'll measure how often it agrees with human reviewers, check it for bias, and watch it for drift over time.
  • Treat test suites as products. That means sealed "never-trained-on" reference sets, manufactured hard cases, a written pass/fail agreement for each agent capability, and tripwires that catch quiet performance decay.
  • Connect lab results to real life. Show whether pre-release scores actually predict production behavior, and route production failures back into the test sets.
  • Attack the agents. Probe for hidden instructions smuggled in through tool outputs, privilege creep, data leakage, and agents gaming their own scoring.
Must have:
  • BS/MS in CS, engineering, statistics, or similar
  • 7+ years of software engineering on complex, enterprise-grade systems
  • 2+ years as a technical lead or architect
  • Strong Python, plus working experience with at least one framework for orchestrating LLM-based agents
  • You've shipped verification infrastructure for long-running, tool-using agents where the results actually decided release outcomes
  • Deep experience automating test or ML-operations pipelines at scale (containers, orchestration, continuous delivery)
  • Solid grasp of LLMs, retrieval-based systems, and agent architectures
Nice to have:
  • Regulated-industry background (audit, accounting, fintech, regtech)
  • Building systems where decisions must be explainable and auditable
  • Security and privacy compliance exposure
  • Testing systems with multiple cooperating agents, especially where handoffs fail
  • Hands-on use of public agent benchmarks or evaluation tooling (tell us which, and what you'd change)
Immediate interview and start - send resume today for immediate consideration!