Infrastructure / Site Reliability Engineer (SRE)

Talent network
Remote
Early applicant

Mercor logo
Posted by Mercor

Mercor connects exceptional technical talent with leading organisations working on ambitious technology and AI initiatives. We are looking for experienced Infrastructure / Site Reliability Engineers (SREs) to join a full-time engagement focused on building and operating complex, enterprise-grade infrastructure. We are seeking engineers with strong hands-on experience building, operating, debugging, and scaling sophisticated production systems. The ideal candidate has worked extensively with Kubernetes, AWS, observability platforms such as Datadog, and modern infrastructure tooling.

This is a full-time opportunity, and candidates must be able to commit to full-time engagement.

What You'll Do

  • Build, operate, and improve highly available and scalable production infrastructure.

  • Manage and optimise Kubernetes-based production environments.

  • Design and maintain cloud infrastructure, primarily across AWS.

  • Improve system reliability, availability, scalability, and operational efficiency.

  • Build and maintain observability across infrastructure and applications using Datadog or similar platforms.

  • Investigate production incidents, perform root-cause analysis, and implement durable fixes.

  • Improve monitoring, alerting, logging, tracing, and overall production visibility.

  • Develop automation and internal tooling to reduce manual operational work.

  • Partner closely with software engineering teams on deployments, infrastructure, and production reliability.

  • Contribute to infrastructure architecture and technical decisions for complex distributed systems.

Ideal Background

  • Professional experience in Infrastructure Engineering, Site Reliability Engineering (SRE), Platform Engineering, DevOps, or Production Engineering.

  • Hands-on experience operating complex, enterprise-grade production systems.

  • Strong production experience with Kubernetes.

  • Strong experience with AWS and cloud-native infrastructure.

  • Experience with Datadog, Prometheus, Grafana, or comparable observability platforms.

  • Experience with Infrastructure as Code using Terraform, Pulumi, or equivalent technologies.

  • Strong understanding of distributed systems, networking, containers, Linux, and cloud architecture.

  • Experience building or maintaining CI/CD and production deployment infrastructure.

  • Strong debugging, troubleshooting, and incident-response capabilities.

  • Proficiency in at least one programming or scripting language, such as Python, Go, or Bash.

Strong Signals

  • Experience operating Kubernetes and cloud infrastructure at significant production scale.

  • Experience supporting high-traffic or mission-critical applications.

  • Experience building infrastructure or platform tooling used by large engineering organisations.

  • Ownership of production reliability, on-call operations, incident response, or capacity planning.

  • Experience working within sophisticated, large-scale distributed systems.

  • Demonstrated improvements to SLOs/SLIs, observability, deployment reliability, infrastructure performance, or operational efficiency.

Why Join

  • Solve challenging reliability, scalability, and performance problems across enterprise-grade production systems.

  • Work extensively with technologies such as Kubernetes, AWS, Datadog, Terraform/Pulumi, and modern cloud-native tooling.

  • Take meaningful ownership of production reliability, observability, infrastructure architecture, and operational improvements.

  • Competitive hourly compensation reflecting your experience and technical expertise.

  • Join a network of highly skilled engineers working on ambitious projects with leading technology and AI organisations.

We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

Contract and Payment Terms

  • You will be engaged as an independent contractor.
  • This is a fully remote role that can be completed on your own schedule.
  • Projects can be extended, shortened, or concluded early depending on needs and performance.
  • Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.
  • Payments are weekly on Stripe or Wise based on services rendered.
  • Please note: We are unable to support H1-B or STEM OPT candidates at this time.

About Mercor

Mercor partners with leading AI labs and enterprises to train frontier models using human expertise. You will work on projects that focus on training and enhancing AI systems. You will be paid competitively, collaborate with leading researchers, and help shape the next generation of AI systems in your area of expertise.

Earn rewards by referring

Share the referral link below and earn 20% of your referral's earnings across all their projects on Mercor. Your Earnings Cap is 4× their hourly rate on their first accepted role. Referral limits apply. See the Referral Policy for current limits.
Referral limits apply. See the Referral Policy for current limits.
Don't know who to refer? Find relevant LinkedIn connections here.


One interview, real results
AI experts share how Mercor made hiring faster, fairer, and easier — with just one interview.

Posted 4 hours ago

$200 / hr

Part-time · Remote