Some imagine the future. At SEGULA, we build it.
Misiune
- Coordinate and drive Site Reliability Engineering (SRE) and Cloud Operations activities with engineering and other stakeholders.
- Ensure stable, secure, and reliable operation of cloud-based applications and services across platforms such as AWS and Alibaba Cloud.
- Take ownership of P1/P2 incidents, lead Root Cause Analysis (RCA), and implement sustainable solutions.
- Improve system reliability through automation, monitoring/observability, performance optimization, and proactive risk management.
- Define and monitor SLIs, SLOs, SLAs, and operational KPIs to ensure service quality and continuous improvement.
- Plan and coordinate releases, platform updates, Kubernetes deployments, databases, and middleware changes.
- Coordinate integration, regression, and acceptance testing to ensure stable and production-ready deployments.
- Support NIS2 compliance in areas such as operational security, monitoring, incident management, and risk management.
- Collaborate with development, platform, architecture, security, and customer teams to ensure efficient delivery and operations.
- Drive improvements in automation, scalability, stability, and operational processes.
- Support service planning and ensure operations meet customer and contractual commitments.
- Mentor SRE team members and share technical expertise and best practices.
Profil
- 5–8+ years of experience in cloud application operations, Site Reliability Engineering (SRE), DevOps, or IT operations.
- Proven experience in coordinating complex technical topics or deliveries without direct people-management responsibility.
- Strong experience operating cloud-based applications on at least one major cloud platform (e.g., AWS, Alibaba Cloud).
- Solid knowledge of Linux-based environments and container technologies (Docker, Kubernetes).
- Hands-on experience with incident, problem, and change management in production environments.
- Experience with application updates, release management, and lifecycle processes.
- Understanding of testing in an operations context (integration testing, system validation, release verification).
- Familiarity with test automation and Continuous Integration and Continuous Delivery (CI/CD) pipelines.
- Knowledge of security and compliance practices, ideally in the context of NIS2, ISO 27001, or similar frameworks.
- Experience with monitoring, observability, reliability engineering practices and automation (e.g., Python, Bash, CI/CD).
- Understanding of SRE concepts such as service reliability, operational metrics, SLIs, SLOs, and automation-first operations.
- Structured and solution-oriented working style, with the ability to connect technical and organizational aspects.
- Strong communication and stakeholder management skills.
- Ability to mentor engineers and drive technical improvements through influence and expertise.
- Fluent English required.
Competențe
Cloud Operations
DevOps
Service Reliability
Site Reliability Engineering
