Staff Software Engineer, Code RL

Anthropic · San Francisco, CA | New York City, NY | Seattle, WA

stafffull timeonsite
Apply on Anthropic’s site →

Role at a Glance

Build the engineering backbone for Claude's code-generation capabilities. As a Staff Software Engineer on Anthropic's Code RL team, you'll rotate across research groups, design frameworks and APIs they run experiments on, and co-own the reliability of production RL systems. Locations: San Francisco, New York City, or Seattle. Compensation: $405,000–$625,000/yr. Hybrid schedule. Visa sponsorship available.

What This Role Is About

This is a platform-engineering role embedded inside AI research. The Code RL team owns the reinforcement learning infrastructure that shapes how Claude learns to write, understand, and reason about code. Rather than conducting RL research yourself, you'll build the substrate that makes research faster and safer: move into a research team, map their engineering bottlenecks, design the tooling and APIs they need, transfer clean ownership, then rotate again. In parallel, you'll hold a standing stake in production RL system health — monitoring, regression detection, and triage tooling included.

About Anthropic

Anthropic is an AI safety company whose defining bet is that building AI responsibly and building it well are the same goal. Rather than pursuing a wide portfolio of projects, the team concentrates on a small number of high-priority research directions, treating AI development as an empirical science with the same rigor as physics or biology. Cross-disciplinary collaboration — between researchers, engineers, policy staff, and business leaders — is central to how the company operates. Anthropic is a public benefit corporation headquartered in San Francisco, with additional offices in New York City and Seattle.

What You'll Do

  • Architect widely-shared APIs, frameworks, and abstractions with clear interfaces and principled defaults that both researchers and engineers can reliably build on
  • Rotate into research teams for defined periods: learn their systems, identify the right abstractions, build the tooling that accelerates their work, and transfer full ownership before moving on
  • Work hands-on in active research codebases, improving their structure and correctness incrementally without disrupting in-flight experiments
  • Proactively identify silent failure modes and eliminate them through deliberate type safety, well-designed invariants, and targeted test coverage
  • Support production RL systems through contributions to monitoring pipelines, regression detection, and triage tooling so issues surface before they compound
  • Help establish the engineering culture of a new team: code standards, review practices, design patterns, and mentorship for researchers growing their engineering craft

Must-Have Qualifications

  • Expert-level Python: static typing, safe async and concurrency patterns, and performance-conscious code
  • Proven track record designing APIs or frameworks that other teams adopted and continued building on
  • Ability to work effectively in large, fast-moving, or research-style codebases you did not originally write
  • Demonstrated history of anticipating silent failure modes and closing them through architecture, typing, and testing — not after-the-fact patching
  • Clear written and verbal communication across a range of engineering backgrounds, including the ability to document complex system designs for varied audiences
  • Comfort operating in loosely scoped, ambiguous environments and driving problems through to maintainable, durable outcomes

Preferred Qualifications

  • Prior experience building tooling or infrastructure specifically for ML research or reinforcement learning workflows
  • Working knowledge of RL concepts, agentic system design, or LLM training pipelines
  • Background in building or operating large-scale distributed systems
  • Experience building client libraries or SDKs on top of sandboxed, containerized, or remote execution environments
  • Hands-on experience with large-scale data pipelines or dataset lifecycle management
  • History of designing extensible plugin systems or class hierarchies adopted across an organization
  • Experience embedding in or consulting for external teams with successful system handoffs
  • Defined code standards, lint rules, or static analysis approaches rolled out across multiple teams
  • Prior technical lead experience or a track record of setting engineering standards at team level
  • Maintained an open source project

Skills Breakdown: Required vs. Preferred

The non-negotiables are advanced Python (typing, async, performance) and a demonstrated ability to design APIs others build on — plus communication strong enough to document those designs for researchers and engineers alike. Preferred skills span RL domain knowledge, distributed systems, and sandboxed execution platforms. The job description is explicit that no single candidate is expected to hold all of them. This is a Staff-level role: Anthropic expects you to scope ambiguous problems independently, generate cross-team impact, and raise the technical bar for the engineers and researchers around you.

Compensation

Anthropic publishes an annual base salary range of $405,000–$625,000 USD for this position. The band is wide because Staff-level placement varies with depth of experience, domain fit, and internal equity benchmarks. This is base salary only — not on-target earnings — as this is not a sales-compensation role. The company describes its overall package as competitive and references equity alongside other benefits, though specific equity terms are not published in this posting.

Benefits

  • Competitive compensation and benefits package
  • Optional equity donation matching
  • Generous vacation and parental leave
  • Flexible working hours
  • Collaborative office spaces in San Francisco, New York City, and Seattle
  • Visa sponsorship with in-house immigration legal support

Location & Work Model

Offices in San Francisco (HQ), New York City, and Seattle — candidates must be based at or willing to relocate to one of these cities. Anthropic uses a hybrid model requiring on-site presence at least 25% of the time; this role may require more depending on team needs. Visa sponsorship is available: Anthropic retains an immigration attorney to assist, though approval is not guaranteed for every candidate or role.

Frequently asked questions

Does this role require prior reinforcement learning experience?

RL familiarity is preferred, not required. The minimum qualifications center on Python depth, API and framework design, and comfort in research-style codebases. The posting is explicit that a strong platform engineer who can learn the domain on the job is still a fit.

Is this a remote-friendly position?

No. Anthropic uses a hybrid on-site policy requiring staff to be present in one of the three offices — San Francisco, New York City, or Seattle — at least 25% of the time. Some roles, which may include this one, require more on-site presence than that minimum.

Does Anthropic sponsor work visas for this position?

Yes. Anthropic offers visa sponsorship and retains an immigration attorney to support the process. Sponsorship is not guaranteed for every candidate or role, but the company states it will make every reasonable effort to obtain a visa for candidates who receive an offer.

What does 'embedding with research teams' look like day-to-day?

Rather than permanently owning a single system, you move into a research team for a defined stretch, learn their engineering needs, build the frameworks or APIs that accelerate their work, and transfer full ownership before rotating to the next team. Alongside rotations, you hold a standing responsibility for production RL system health.

Is a formal degree required to apply?

Anthropic lists a Bachelor's degree as the minimum education level but explicitly accepts equivalent combinations of education, training, and professional experience in a relevant field — so candidates without a traditional four-year degree are encouraged to apply.

Are you a fit for this role?

See how your background matches Staff Software Engineer, Code RL — and the 20 live jobs that fit you best. No signup, no email, nothing stored.