Founding AI Research Lead - Agentic AI Lab
fabrion
San Francisco Bay Area
Posted Sep 6, 2026
- Full-time
- Engineering
Job description
**Founding AI Research Lead - Agentic AI Lab** *San Francisco Bay Area | Full time* Backed by 8VC, we are building a world-class team to tackle one of industry's most critical problems: trusted AI for enterprise operations. **About the Role** Fabrion is designing the future of enterprise AI infrastructure, grounded in agents, knowledge graphs, and multi-tenant governance. We are working on research inside the Agentic AI Lab to train and evaluate specialized models for mission-critical enterprise work. The direction is specific and ambitious. We share the full thesis under NDA during the interview process. What we can say here: the program has committed design partners with production data access, dedicated compute, a benchmark-first plan with clear go and no-go gates, and a platform team that has already built the governance and serving layer your models will run behind. This is full-cycle research: problem formulation, data, training, evaluation, and deployment, with your name on the results. **Core Responsibilities** - Own the research agenda: model and training design, evaluation protocol, and the publication plan - Take models from public benchmark results to live customer shadow deployments, with gates you define and defend - Set the benchmark discipline: strong baselines first, published comparables cited, results that survive scrutiny - Lead and grow a small team (ML engineer, data engineer, contractors) and pair closely with the founders and platform team - Write technical plans internally and papers externally when results warrant it **Desired Experience** - Hands-on experience training sequence models, owning the tokenizer, the training loop, and the evaluation, not only fine-tuning through APIs - Strong background in at least two of: reinforcement learning (especially offline and imitation settings), sequence decision modeling, structured or constrained generation, learning from event and log data - A track record of shipping research into a product or landing a rigorous benchmark result - PhD in machine learning or a closely related field, or an equivalent research record - Preferred Tech Stack - PyTorch, the Hugging Face ecosystem, experiment tracking and reproducible training pipelines, modern cloud data warehouses, evaluation harness engineering **Soft Skills & Mindset** - Comfortable as the most senior researcher in the room: setting direction under ambiguity and writing decisions down - Rigor over hype: you distrust your own results until the baselines agree - A teacher's instinct: part of this role is turning strong engineers into researchers **Why This Role Matters** We believe specialized models built on governed enterprise data can run real, multi-billion-dollar workflows. Your work will not be buried in research reports. It will be benchmarked in public, deployed to real customers, and activated by hundreds of thousands of decisions.