Senior Site Reliability Engineer (m/f/d)

  • CDI
  • Temps plein
  • Au moins 5 ans d'expérience
  • BUT, Licence, Bac+3
  • ADMINISTRATEUR SYSTEME

Some imagine the future. At SEGULA, we build it.


Mission

  • Coordinate and drive Site Reliability Engineering (SRE) and Cloud Operations activities with engineering and other stakeholders.
  • Ensure stable, secure, and reliable operation of cloud-based applications and services across platforms such as AWS and Alibaba Cloud.
  • Take ownership of P1/P2 incidents, lead Root Cause Analysis (RCA), and implement sustainable solutions.
  • Improve system reliability through automation, monitoring/observability, performance optimization, and proactive risk management.
  • Define and monitor SLIs, SLOs, SLAs, and operational KPIs to ensure service quality and continuous improvement.
  • Plan and coordinate releases, platform updates, Kubernetes deployments, databases, and middleware changes.
  • Coordinate integration, regression, and acceptance testing to ensure stable and production-ready deployments.
  • Support NIS2 compliance in areas such as operational security, monitoring, incident management, and risk management.
  • Collaborate with development, platform, architecture, security, and customer teams to ensure efficient delivery and operations.
  • Drive improvements in automation, scalability, stability, and operational processes.
  • Support service planning and ensure operations meet customer and contractual commitments.
  • Mentor SRE team members and share technical expertise and best practices.

Profil

  • 5–8+ years of experience in cloud application operations, Site Reliability Engineering (SRE), DevOps, or IT operations.
  • Proven experience in coordinating complex technical topics or deliveries without direct people-management responsibility.
  • Strong experience operating cloud-based applications on at least one major cloud platform (e.g., AWS, Alibaba Cloud).
  • Solid knowledge of Linux-based environments and container technologies (Docker, Kubernetes).
  • Hands-on experience with incident, problem, and change management in production environments.
  • Experience with application updates, release management, and lifecycle processes.
  • Understanding of testing in an operations context (integration testing, system validation, release verification).
  • Familiarity with test automation and Continuous Integration and Continuous Delivery (CI/CD) pipelines.
  • Knowledge of security and compliance practices, ideally in the context of NIS2, ISO 27001, or similar frameworks.
  • Experience with monitoring, observability, reliability engineering practices and automation (e.g., Python, Bash, CI/CD).
  • Understanding of SRE concepts such as service reliability, operational metrics, SLIs, SLOs, and automation-first operations.
  • Structured and solution-oriented working style, with the ability to connect technical and organizational aspects.
  • Strong communication and stakeholder management skills.
  • Ability to mentor engineers and drive technical improvements through influence and expertise.
  • Fluent English required.

Compétences

Cloud Operations
DevOps
Service Reliability
Site Reliability Engineering