HomeSearchAi Engineer Jobs › Staff AI Engineer - Agent Architecture & Behavior

Staff AI Engineer - Agent Architecture & Behavior

Artisan

San Francisco, CA, United States · staff
$250,000–$325,000 USD
More Ai Engineer jobs: Ai Engineer jobsAi Engineer salary

Job at a glance

$250,000–$325,000 USD
Salary
San Francisco, CA, United States
Location
Staff
Seniority
Artisan
Employer
Ai Engineer jobs
Category

Artisan · Full-time · In person in San Francisco Base salary: $250,000–$325,000 USD annually. Equity: 0.15%–0.30%. US visa sponsorship available. Build something new at the frontier of applied AI At Artisan, we're working on a new, ambitious project that will push the boundaries of what agentic AI can do. We're keeping the product details private ahead of launch, but we can tell you this: the technical problems are substantial, the scope for invention is real, and this hire will shape the core technology.

We're looking for a hands-on technical lead to design and build the underlying AI architecture. You'll work across agent behavior, complex multi-agent systems, tool use, context, and evaluation, taking promising ideas through to dependable production software. This is an individual-contributor role with broad technical ownership. You'll make consequential architecture decisions, write the hardest parts of the system, and work closely with our existing engineers and leadership.

You should enjoy both exploring an uncertain problem and doing the detailed engineering required to make a solution work. What you'll own Agent architecture and behavior. Design and implement agent execution loops, planning strategies, tool interfaces, and verification. Turn ambiguous technical requirements into clear system boundaries and working code. Multi-agent systems. Build delegation, coordination, context sharing, and result synthesis.

Handle concurrent work, conflicting updates, cancellation, and stale results. Establish when a multi-agent approach improves on a simpler baseline. Reliable execution. Make complex, stateful workflows resilient to interruptions and partial failures. Build checkpoints, recovery strategies, and appropriate human intervention into the architecture. Context, memory, and reusable methods. Improve retrieval, context construction, persistent state, and skill representation.

Investigate how systems can use feedback and prior experience to perform better without introducing regressions. Evaluation and experimentation. Build realistic evaluations, analyze task trajectories, and turn observed failures into measurable improvements. Compare approaches using quality, reliability, latency, and cost. Model and tooling decisions. Evaluate models and emerging techniques, prototype promising approaches, and make informed build-versus-buy decisions.

Choose tools because they solve the problem, and be willing to replace them when the evidence changes. Technical leadership. Set engineering standards, review important design decisions, and help the team implement a coherent AI system. Stay close to the product and accountable for what ships. You'll partner with product and infrastructure engineers on production services, integrations, secure execution, and observability.

You'll own the AI architecture and its effectiveness, with implementation shared across the team. What we're looking for You have personally built and shipped a substantial agentic system. Production use or rigorous, reproducible open-source work matters more than the name of a framework or employer. You have deep practical experience with LLM tool use, planning, context engineering, and evaluations.

You have implemented multi-agent coordination or substantial parallel agent/tool execution and can explain its failure modes. You have hands-on experience with browser or computer automation in an agentic system, including observing state, verifying effects, and recovering when an interface or execution path fails. You are an excellent software engineer in Python, TypeScript, or a comparable language.

You are comfortable with asynchronous services, state machines, persistence, concurrency, retries, and cancellation. You know which decisions belong to a model and which guarantees must be enforced in code. You can reason carefully about permissions, untrusted inputs, uncertain external outcomes, and human approvals. You can design meaningful experiments, debug real system be

Search all live jobs — free, no account →