Intellibus Ventures LLC

Lead Site Reliability Engineer (SRE) – DevOps & Observability

Atlanta, GA, US$156,000-$176,800Posted 2 days ago

Job Description

Imagine working at Intellibus to engineer platforms that impact billions of lives around the world. With your passion and focus, we will accomplish great things together.

Our Platform Engineering Team is looking for experienced DevOps / SRE leaders who can help build highly reliable, observable, and automated infrastructure supporting mission-critical applications.

We are looking for hands-on technical leaders with deep experience in Datadog, infrastructure automation, configuration management, Terraform/Chef, Bash/Shell scripting, cloud infrastructure, and production systems.

The ideal candidate will also have a strong understanding of Java-based applications and Java coding, as this role will work closely with Java engineering teams and distributed application platforms.

We are looking for Architects who can do the below but not limited to

  • Lead DevOps/SRE initiatives across mission-critical environments.
  • Design and implement observability solutions using Datadog.
  • Build dashboards, monitors, alerts, logs, traces, and actionable operational metrics.
  • Define and improve SLIs, SLOs, SLAs, and error-budget practices.
  • Automate infrastructure provisioning and configuration using Terraform, Chef, or similar tools.
  • Develop and maintain Bash/Shell scripts for infrastructure and operational automation.
  • Manage deployment, configuration, and environment automation across development, QA, and production.
  • Troubleshoot complex production, infrastructure, networking, and application issues.
  • Improve system availability, scalability, performance, and reliability.
  • Support CI/CD pipelines and automated deployments.
  • Work closely with Java engineering teams to understand application behavior, performance, dependencies, and production issues.
  • Analyze Java applications from an operational perspective, including JVM behavior, memory, CPU, threads, logs, and application performance.
  • Participate in incident response, root-cause analysis, and pos

Apply for this role

Keep looking

Similar Remote IT And Developer Jobs