About the Role
We are looking for a hands-on DevOps Engineer with 3–4 years of production experience to independently own and operate a client’s cloud and DevOps environment.
You will own the day-to-day reliability, deployments, automation, monitoring, troubleshooting, and continuous improvement of the client's production infrastructure across AWS, Azure, and/or GCP.
This client-facing role requires strong technical communication, ownership, and the ability to solve operational problems independently. You will also guide and coordinate 1–2 engineers working with you on the client environment.
You should be comfortable working with Kubernetes, CI/CD, Infrastructure as Code, Linux, automation, observability, security, and production incident management.
Key Responsibilities:
What You Will Own
- End-to-end technical ownership of one client environment
- Day-to-day health, reliability, and operational excellence of the environment
- Deployments, infrastructure changes, and production support
- Monitoring, alerting, and incident response
- Infrastructure automation and continuous improvement
- Client technical communication and requirement understanding
- Coordination, guidance, and technical review of 1–2 engineers
- Documentation, reporting, and operational knowledge management
- Identification of opportunities for reliability, security, performance, and cost improvement
Cloud Infrastructure Operations
- Own and operate workloads running on managed Kubernetes platforms such as EKS, AKS and GKE, including deployments, scaling, rolling updates, troubleshooting and cluster maintenance.
- Implement, test and maintain backup, restore and recovery workflows for object storage and managed databases.
- Operate and troubleshoot core cloud services including compute, storage, databases, networking, IAM/RBAC and load-balancing services across AWS, Azure and/or GCP:
- Compute: EC2 / Azure VM / Compute Engine
- Storage: S3 / Azure Blob Storage / Google Cloud Storage
- Databases: RDS / Azure SQL or Azure Database / Cloud SQL
- Networking: VPC / Virtual Network / VPC Network
- Security: IAM / Azure AD & RBAC / GCP IAM
- Load Balancers and Networking Services
CI/CD & Automation
- Maintain CI pipelines using tools such as GitHub Actions, GitLab CI, Jenkins or similar for build, test, and artifact management.
- Manage application delivery pipelines using CD tools such as Spinnaker, ArgoCD, Flux, or similar, implementing deployment strategies like Blue-Green or Canary deployments.
- Troubleshoot failed builds, deployments, and releases and identify opportunities to improve deployment reliability and speed.
- Automate repetitive operational activities and reduce manual intervention and operational toil.
- Write and maintain automation scripts (Bash/Python) to streamline deployment workflows and reduce manual effort.
Security & Compliance
- Implement and maintain infrastructure security practices including vulnerability management, IAM/RBAC, least privilege, secrets management, and OS hardening for Linux workloads.
Monitoring, Logging & Performance
- Configure dashboards, alerts, and log pipelines using observability platforms such as Coralogix, Grafana, New Relic, Datadog, or cloud-native monitoring tools.
- Troubleshoot performance and availability issues across infrastructure, containers, networking, and applications.
- Proactively identify reliability and performance issues before they become client-impacting incidents. Own incident response for assigned environments, contribute to RCA, and implement preventive actions.
Linux & System Administration
- Manage Linux-based workloads across compute nodes and Kubernetes clusters, including patching, tuning, troubleshooting, and log analysis.
- Perform OS-level maintenance, troubleshooting, incident response, and recovery while maintaining production availability and security.
Client & Stakeholder Management
- Act as the primary technical DevOps contact for the assigned client.
- Understand client requirements and translate them into technical actions.
- Communicate infrastructure health, incidents, risks, and planned changes clearly.
- Handle routine technical discussions independently.
- Set appropriate expectations and escalate risks when required.
- Maintain professional and responsive communication during production incidents.
- Build trust through predictable execution and transparent communication.
Team Coordination & Mentoring
- Coordinate, guide, and review the work of 1–2 engineers supporting the client environment.
- Review technical work and ensure adherence to engineering standards.
- Help team members troubleshoot complex issues.
- Share knowledge and improve team documentation.
- Proactively escalate capability, delivery, or performance risks.
Required Skills & Experience
- 3–4 years of hands-on DevOps / Cloud Operations experience.
- Strong hands-on production experience with at least one major cloud platform (AWS, Azure, or GCP), with working knowledge of core services across cloud environments.
- Practical experience with Kubernetes deployments in production (preferably EKS).
- Good understanding of GitHub (SCM) and building CI workflows in GitHub Actions.
- Hands-on experience with Infrastructure as Code and automation using Terraform, CloudFormation, Bash and/or Python.
- Experience administering Linux environments (Ubuntu / Amazon Linux).
- Working knowledge of monitoring/logging tools – Cora Logix or similar (New Relic, Grafana, CloudWatch).
What We Look For
- Strong ownership and accountability
- Ability to independently troubleshoot production issues
- Ability to own one client environment end-to-end
- Strong client-facing communication
- Ability to coordinate and guide 1–2 engineers
- Practical and structured problem-solving
- Willingness to learn and adapt to new technologies
- Strong focus on reliability, security, automation, and operational excellence
Success in This Role
After becoming fully effective, you should be able to:
- Independently own and operate one client environment
- Resolve most production issues without constant supervision
- Maintain reliable deployments and infrastructure
- Proactively identify and fix operational risks
- Communicate effectively with the client
- Coordinate 1–2 engineers
- Continuously improve automation, reliability, and cost efficiency