Machine Learning Engineer, Applied AI
About this role
Our Mission
Rebuild how the world works, to make institutions work better for the people they serve.
About Brain Co.
Brain Co. builds AI-native operating systems for large, regulated institutions. Each system is built for a specific industry, powered by agents that push real workflows forward. Underneath it all is Atlas, our proprietary platform that keeps customers in control, secure by design, and never locked into one model.
Why Now
Brain Co. is entering its next phase of production deployments on a national scale with an elite team built from Palantir, Google, Meta, and Nvidia, and a growing footprint across government, insurance, health, and financial services.
Joining now means shaping both the company and a new category of applied AI. Every project here ships to production and is expected to create measurable customer value and impact.
You'll work alongside exceptional peers on some of the hardest problems in applied AI. It’s the kind of work you'll still be proud of in ten years from now.
Machine Learning Engineer, Applied AI
About the Role
So much of the work society depends on is still slower and harder than it should be. Permits take months. Claims sit unresolved. And AI hasn't changed that — because the bottleneck isn't the models. It's the institutional context AI needs to do the work: rules, history, relationships, and judgment scattered across people, documents, and legacy systems.
BrainCo exists to fix that. We build agent-native operating systems for the institutions society depends on, and our products are the first of their kind in the world — we were the first, anywhere, to fully automate construction permitting, and we're now doing the same across insurance and other industries. There is no playbook here, because no one has built this before.
As a Machine Learning Engineer on Applied AI, your work begins where the demo ends: getting a model to look impressive is the easy part; making it a production decision system an institution stakes its process on is the job. The problems come in every shape — custom vision model pipelines that check blueprints against building codes at 95%+ accuracy, agents that untangle policy stacks to reveal coverage gaps, systems that predict from clinical records whether a patient is on their care path — and you'll own them end-to-end, from ambiguous customer problem to the eval that catches a whole class of errors.
This is frontier ML applied where it's hardest and matters most. The problems are underspecified, the documents are brutal, the accuracy bar is institutional-grade — and the feedback loops are real, because our systems move real workflows forward every day.
Who We're Looking For
You understand how machine learning actually works — not just the tooling, but the philosophy underneath: what a loss function really optimizes, how generalization breaks under distribution shift, why evaluation is where systems quietly go wrong. And you live at the bleeding edge of modern AI, with hard-won instincts for squeezing the most out of LLMs and agentic systems — prompting, fine-tuning, tool use, and reasoning. That combination is the job: you know when a fine-tuned segmentation model beats a VLM, when a rule engine beats both, and how to compose all three into a system more accurate than any single model. You treat frontier models as components to be measured, pushed, and engineered — never as magic.
Most of all, you're energized by building things that have never existed, and comfortable when the problem, the data, and the definition of success all have to be invented at once.
The Problems You'll Work On
Composite AI systems and credit assignment. Our most demanding systems chain vision transformers, segmentation models, VLM reasoning, and rule engines. When the pipeline is wrong, which component failed? One of the most interesting open problems in applied ML.
Document understanding beyond the frontier. Blueprints, site plans, policy stacks, contracts, clinical records — dense, multimodal documents that break off-the-shelf models. You'll build models that actually read them.
Agents that learn from real work. Our deployments generate verified, ground-truth outcomes on every decision — reward signals most labs can only simulate. You'll help design the data, evals, and training loops to build and fine-tune agents on them.
Evaluation as a product discipline. When a regulator has to trust your system, evals are the product. You'll build eval suites and failure-mode taxonomies rigorous enough to earn institutional sign-off.
Institutional Intelligence that compounds. Every verified correction improves the system twice: the corrected fact percolates to every application, and the system that builds the intelligence learns to build it better. You'll work on both loops.
In This Role, You Will:
Turn ambiguity into shipped systems — from no problem statement, no labeled data, and no agreed definition of success, to well-posed ML problems and production deployments.
Own AI systems end-to-end. There is no handoff: the person who trains the model owns its behavior in production.
Work at the research frontier with production stakes, applying LLMs, RL fine-tuning, and agentic systems where the output is a decision an institution acts on.
Work directly with the institutions we serve — permit reviewers, underwriters, compliance officers — to understand how decisions actually get made and ensure your systems change how the work gets done.
Engineer for production reality, navigating accuracy, latency, cost, and reliability in environments far messier than any benchmark.
Raise the bar across the company through design reviews, our internal paper club, and the shared playbook for AI systems institutions can trust.
