Site Reliability Engineer

Published 8 of September

Collaborate closely with internal and external stakeholders, developers, and managers to ensure timely deliverables.

Develop, maintain, and scale configuration management and automated provisioning using Ansible across cloud virtual machines and hybrid environments.

Build, secure, and operate scalable infrastructure on public cloud platforms (predominantly Azure), managing compute instances, networking, and storage.

Partner tightly with the security team to identify vulnerabilities, execute mitigation strategies, install security components like EDR, and ensure compliance standards are met across all managed systems.

Administer and optimize identity management systems, including Azure Active Directory / EntraID, user permissions, and SSO configurations, ensuring secure and seamless access.

Write, review, and refactor IaC modules using Terraform within a structured Git repository model.

Automate provisioning, configuration and operational recovery with a reliability-first mindset

Integrate robust monitoring and logging stacks to ensure high availability, troubleshoot complex infrastructure incidents, and participate in continuous service reliability operations.

Maintain cost-awareness and resource tagging policies to support platform financial governance (FinOps).

Participate in the 24/7 rotational hotline to maintain system resilience.


Minimum requirements

Solid background as a Site Reliability Engineer, DevOps Engineer, or Systems Engineer managing production environments in the cloud.

Strong operational experience with public cloud platforms, particularly Microsoft Azure.

Experience with identity management frameworks (Azure Active Directory / EntraID, SSO, user permissions) and implementing security tools (EDR, vulnerability scanners).

Proficiency with Terraform for state management and modular infrastructure deployments and Ansible for automated configuration, patching, and server management.

Strong autonomy and critical thinking skills, with a focus on collaboration and ability to work in cross-functional teams.

Excellent communication skills (written, listening, and speaking), a collaborative mindset, and a proactive attitude.

Fluency in English (mandatory).

Additional/Preferable Skills

Experience with container orchestration platforms (Kubernetes / AKS).

Proficiency in CI/CD pipeline management using GitLab and GitLab CI/CD.

Familiarity with the Atlassian suite, particularly JSM (Jira Service Management) for incident, request, and workflow tracking.

Background in observability stacks (Prometheus, Grafana, Loki).

Relevant certifications such as Certified Azure Administrator (AZ-104).