Roadie Logo

Roadie

Lead Site Reliability Engineer

Sorry, this job was removed Sorry, this job was removed at 05:06 p.m. (PST) on Tuesday, Apr 29, 2025
Remote
Remote

Similar Jobs

13 Days Ago
Remote or Hybrid
Urbandale, IA, USA
123K-184K Annually
Senior level
123K-184K Annually
Senior level
Artificial Intelligence • Cloud • Internet of Things • Machine Learning • Analytics • Industrial
The Lead SRE-React Front-End ensures high reliability, security, and performance of web and mobile solutions, provides mentoring, and manages complex problem resolution.
Top Skills: Aws CloudfrontJavaScriptNode.jsReact
3 Days Ago
Remote or Hybrid
Boston, MA, USA
Senior level
Senior level
Artificial Intelligence • Big Data • Information Technology • Software
Lead Site Reliability Engineer responsible for high-performance cloud platform management, driving SRE processes, team leadership, and ensuring FedRAMP compliance.
Top Skills: AnsibleAWSAzureBashCloudFormationCrossplaneDockerGCPGitGitlabGoJenkinsKubernetesPythonTerraform
13 Days Ago
Remote or Hybrid
New York, NY, USA
120K-160K Annually
Senior level
120K-160K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Lead Site Reliability Engineering efforts focusing on CMDB and asset data quality. Design integrations, manage documentation, and ensure data accuracy for operations.
Top Skills: AWSAzureDevice 42EracentGCPServicenowUcmdb

Roadie, a UPS company, is a leading logistics and delivery platform that helps businesses tackle the complexities of modern retail with unmatched delivery coverage, flexibility and visibility. Reaching 97% of U.S. households across more than 30,000 zip codes — from urban hubs to rural communities — Roadie provides seamless, scalable solutions that meet a variety of delivery needs. 

With a network of more than 310,000 independent drivers nationwide, Roadie offers flexible delivery solutions that make complex logistics challenges easy, including solutions for local same-day delivery, delivery of big and bulky items, ship-from-store and DC-to-door. 

Roadie is seeking a Lead Site Reliability Engineer to join our growing Technical Operations Team. We are looking for a leader with a proven track record of managing high-performing SRE teams in high-availability, mission-critical environments. The ideal candidate is a strategic problem solver with deep expertise in site reliability best practices, DevOps principles, AWS and GCP, Kubernetes, and automation. You will play a key role in driving reliability, scalability, and operational excellence across our platform.

What You'll Do

  • Lead and mentor teams focused on enhancing platform reliability, optimizing uptime, and improving software delivery, observability, and infrastructure operation
  • Architect, maintain, and optimize production and non-production Kubernetes clusters (EKS), as well as Elasticsearch (ES), MSK, RDS, and ElastiCache (Redis) clusters
  • Design, deploy, and manage monitoring and logging solutions using Prometheus, Loki, Thanos, Grafana, OpenTelemetry, and New Relic
  • Strategize and collaborate with cross-functional teams to proactively identify bottlenecks, optimize resource utilization, and prevent system failures
  • Define and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability improvements
  • Automate and streamline operational tasks, reducing toil and increasing efficiency across engineering teams
  • Plan and forecast service capacity and demand, optimize costs, and fine-tune system performance
  • Lead troubleshooting initiatives, post mortems, and resolve production and non-production incidents, ensuring high availability and performance
  • Participate in and manage a 24/7 on-call rotation, responding to incidents and driving post-mortem improvements
  • Willingness to work non-standard hours to facilitate production upgrades or deployments on occasion

Technology We're Using Now

  • Python, Ruby on Rails, Golang
  • React/Redux, Objective-C and Swift, Android
  • Postgres, Redshift, Redis, Kafka
  • AWS/GCP
  • Docker/Kubernetes
  • OpenTelemetry/Prometheus/Thanos/Loki/Grafana/New Relic/Sentry
  • Git/CircleCI
  • ArgoCD

What You Bring

  • 6+ Years in various SRE roles
  • 6+ Years in various DevOPS/System Engineering roles
  • 3+ Years in leading and managing SRE teams
  • 6+ Years of experience building and managing production Kubernetes infrastructure
  • 7+ Years experience with popular scripting languages (Python, Ruby, Bash, etc.)
  • Experience with Infrastructure as code such as Terraform or Crossplane
  • Experience with CI/CD Development tools (CircleCI, etc.)
  • Experience with GitOPS Tools (ArgoCD)
  • Experience using a broad range of AWS technologies (RDS, ElasticSearch, VPC, EKS, S3, CloudFront, MSK, Elasticache, CloudWatch, etc.)
  • Experience developing and maintaining YAML templating systems (Helm charts, Kustomize, etc)
  • Must be able to work independently, be self-motivated and handle multiple priorities
  • Comfortable working in a fast-paced agile environment

Finally, a willingness to admit what you don’t know, and learn what you need to learn quickly.

Why Roadie? 

  • Competitive compensation packages 
  • 100% covered health insurance premiums for yourself
  • 401k with company match
  • Tuition and student loan repayment assistance (that’s right - Roadie will contribute directly to your existing student loans!) 
  • Flexible work schedule with unlimited PTO 
  • Monthly 3-day weekends
  • Monthly WFH stipend 
  • Paid sabbatical leave- tenured team members are given time to rest, relax, and explore
  • The technology you need to get the job done

This role is not eligible for Visa sponsorship. Applicants must be authorized to work for any employer in the U.S.

What you need to know about the San Francisco Tech Scene

San Francisco and the surrounding Bay Area attracts more startup funding than any other region in the world. Home to Stanford University and UC Berkeley, leading VC firms and several of the world’s most valuable companies, the Bay Area is the place to go for anyone looking to make it big in the tech industry. That said, San Francisco has a lot to offer beyond technology thanks to a thriving art and music scene, excellent food and a short drive to several of the country’s most beautiful recreational areas.

Key Facts About San Francisco Tech

  • Number of Tech Workers: 365,500; 13.9% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Google, Apple, Salesforce, Meta
  • Key Industries: Artificial intelligence, cloud computing, fintech, consumer technology, software
  • Funding Landscape: $50.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Sequoia Capital, Andreessen Horowitz, Bessemer Venture Partners, Greylock Partners, Khosla Ventures, Kleiner Perkins
  • Research Centers and Universities: Stanford University; University of California, Berkeley; University of San Francisco; Santa Clara University; Ames Research Center; Center for AI Safety; California Institute for Regenerative Medicine
By clicking Apply you agree to share your profile information with the hiring company.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account