We use cookies. Find out more about it here. By continuing to browse this site you are agreeing to our use of cookies.
#alert
Back to search results
New

Senior Site Reliability Engineer

Skill
United States, Iowa, Urbandale
Jul 31, 2026
Overview

Placement Type:

Temporary

Salary:

$71.13-76.13 Hourly

Hourly, depends on experience/education

Start Date:

Aug 24, 2026

Senior Software Engineer / Site Reliability Engineer

About the Role

Aquent Studios is looking for a Senior Software Engineer who combines strong software engineering ability with the curiosity and interpersonal skill to dive into unfamiliar code, unfamiliar languages, and unfamiliar teams - and help them get to resolution. This is not a scripting/automation-only role, nor is it a pure feature-development role. Success here means being able to read and reason about production code you didn't write, in languages you may not know well, quickly enough to be useful in the moment, and doing so in a way that builds trust with the engineering teams you support.

Roughly 80% of this job is relationship-building and cross-team collaboration - but that only works if you bring real technical credibility into the room. You'll use observability tools and logs to trace a problem to its source, then sit down with an engineering team to work through the code together toward a fix.

What You'll Do



  • Use observability and monitoring tools (e.g., Datadog, logs, traces) to identify, isolate, and diagnose issues across complex, multi-service systems.
  • Partner directly with engineering teams to debug and resolve issues in codebases and languages you may not have prior exposure to - reading and reasoning about unfamiliar code under time pressure.
  • Act as Incident Commander, coordinating recovery efforts for large or complex system failures.
  • Identify failure modes as you investigate the code and build automated solutions for detection, failure handling, and recovery.
  • Lead cross-team problem-solving for resiliency, reliability, and security issues, identifying and organizing the right resources and acting as a technical advisor.
  • Lead the development of Service Level Objectives (SLOs) across multiple parts of the system.
  • Evaluate and implement design improvements that enhance cost, quality, performance, and security, along with the metrics to track them.
  • Lead post-incident reviews (postmortems), driving out learnings and ensuring changes to automation, recovery processes, and documentation are shared across the SRE team.
  • Collaborate closely with other SREs and product/engineering teams to ensure features and systems meet reliability and business needs.
  • Maintain and produce support and playbook documentation as needed.
  • Stay current on industry and technical innovations relevant to reliability engineering.


What You Bring

Core Orientation



  • Genuine software engineering experience-you've built and shipped features, not just automation scripts-and can read/write production code with confidence.
  • A clear understanding of what SRE is for: you see yourself as someone who supports and strengthens other teams' systems, not primarily as a feature developer.
  • High curiosity and comfort operating with incomplete information - you're willing to jump into an unfamiliar system or codebase and work your way to an answer.
  • Strong interpersonal and communication skills; ability to build trust quickly with engineering teams and lead through ambiguity during incidents.


Cloud & Infrastructure



  • AWS Cloud Services
  • Kubernetes
  • Terraform (infrastructure as code)


Observability & Monitoring



  • Datadog
  • Experience using logs and tracing tools to perform root-cause diagnosis across distributed systems


Programming Languages



  • Proficiency in at least one production language (e.g., Java, Scala, JavaScript, .NET, Go, Python)
  • Demonstrated ability to quickly ramp up in unfamiliar languages/codebases - this matters more than depth across many languages


Reliability Practices



  • Experience leading or actively participating in Incident Command and postmortems
  • Experience defining and maintaining Service Level Objectives (SLOs)
  • Experience building automated failure detection, handling, and recovery solutions
  • Experience instrumenting and analyzing metrics for cost, quality, performance, and security


Please note: Sponsorship is not available for this role. Applicants must be able to work for any US-based employer.

Please note: This position is open to US-based applicants only and is offered on a W2, hourly basis (no C2C, please).

Please note: This is a fully onsite position, located at our client facility in Urbandale, IA. Remote candidates will not be considered.

Applied = 0

(web-77cf7d65c7-rcc7h)