NYC, Full-Time, In-Person
The Situation
Honcho is ripping. 50x developer growth in three months. 40,000 prosumers and developers. 600+ startups applied to our startup program in eight weeks.
We're a product-coupled neolab solving personal identity for the agentic world. The market currently calls what we do memory, but the vision for Honcho is much more ambitious. We build the foundation for models and agents to know you better than you know yourself. The goal is true 1:1 individual alignment for every human extending their cognition and agency with artificial intelligence.
We aim to model any entity: a person, an agent, a brand, a team, an NPC. A representation that predicts what that entity will do, believe, and choose. All without access to ground truth.
That means separating durable identity from transient state, telling whether a representation that updates is improving, and keeping the methods valid for an entity that changes under measurement.
Evals are the instrument. This role owns it.
The Job
You own evaluation of Honcho end to end, and the loop it drives:
Design the evals. You define the scores. You decide what "better" means for a representation of an entity that changes over time, and you keep that definition current. When the product or the methods move, the evals move with them.
Build the pipelines and harnesses. Data in, labels, versions, reruns, judges. The machinery that lets the team ask a new question this week and get an answer this week.
Run them and harvest insights. Read the traces and the results. Find what's actually broken, not what's easy to measure.
Propose the fix. Changes to the system that produce higher fidelity representations, and the measurement that tells us whether they worked.
Set the next question. Find the gap in our evals, scope the next experiment, run it again.
You also own the infrastructure underneath it. Part of that is building agents that run simulations at a scale we can't reach by hand, because the ceiling on our iteration speed is how many entities we can model and measure at once.
Every part of this loop is yours -- the question, the pipeline, the rerun, the writeup. Nothing gets scoped and handed off, and nothing waits on someone else's sprint.
You
We care what you originated and whether other people used it.
- You've built production evals. Designed and built end to end, then used by a group larger than you to measure a real system. Not toys.
- You know when a result is real. And when the eval is leaking, the judge is drifting, or the delta is noise. Skepticism about your own numbers is the job.
- You can build a measurement from nothing. No benchmark to reach for, no labels, no ground truth. You can construct one, argue for why it's valid, and know what it fails to capture.
- You can code, and maintain what you write. Tests, issues, a release cadence. It keeps running after you've moved on to the next thing.
- You can grok a complex system fast. Dropped into an unfamiliar codebase or system topology, you can build an accurate map of it in hours, not weeks.
- You believe in open. Honcho is open source. You think eval methods and methodology should be as well.
- High conviction. This is a mission-driven team in it for the long haul. We want you in for it too.
- NYC, in person--or ready to move.
Technical Requirements
- Trackable research experience
- first or co-first author at NeurIPS, ICML, ICLR, or equivalent venues
- or independent open-source work of comparable quality
- 3+ years in LLM evaluation, agentic optimization, or an adjacent field
- Strong Python
- Experience leveraging agents for development
- Familiarity with common data pipeline tooling (huggingface, hydra, kafka, SQL, etc.)
A plus: work on memory, identity, personalization, or agent evaluation. Experience scoring representations that update over time. Interests in the cognitive sciences--linguistics, neuroscience, philosophy, psychology. Up-to-date in the open-source AI community.
Compensation
- Base: $220K–$300K + equity
- Plus: standard Plastic benefits
Apply
Two things:
- An eval, benchmark, or eval pipeline other people have actually run. Tell us what you got right and what you'd rebuild.
- Your record in whatever form it exists--resume, portfolio, GitHub, Scholar, personal site, etc.
Our current numbers are public at evals.honcho.dev. Come tell us what's wrong with them.
(Back to Working at Plastic)