Vytvořte si profil, aby vás zaměstnavatelé mohli najít, abyste dostávali vhodnější pracovní nabídky a mohli se rychleji ucházet o práci.
  • Hledání práce
  • Oblíbené položky
  • Vytvořte si CV
    Nové
  • Platy
  • Předplatné

Senior Site Reliability Engineer (AI/ML Platform) @ Link Group

20.000 - 27.000 PLN
Plný úvazek

Link Group

Zahraničí
  • Práce na dálku

Odpovědnosti

Původní popisek. 
  1. Build Bulletproof Observability: You will design and implement the "nervous system" for our AI platform. This means going beyond basic monitoring to build comprehensive observability with robust telemetry, insightful dashboards (Grafana), and intelligent alerting (Prometheus). You will define and track SLOs/SLIs to ensure our services meet their promises and drive improvements when they don't.
  2. Automate Everything: Your mantra is "if you have to do it twice, automate it." You will write clean, effective code in Python or Go to eliminate manual toil, create self-healing systems, and build sophisticated tooling that accelerates incident response and makes deployments safer.
  3. Own the Incident Response Lifecycle: When critical systems fail, you will be on the front lines. You will lead the charge in incident management, participate in a blameless on-call rotation, and conduct insightful post-mortems that lead to real, lasting improvements. You'll build the runbooks that others will rely on.
  4. Engineer a World-Class Deployment Pipeline: You will be a key contributor to our CI/CD ecosystem, building rock-solid integrations, automated safety checks, and seamless rollback capabilities to ensure that we can innovate at speed without sacrificing stability.
  5. Act as a Reliability Partner for Product Teams: You will work side-by-side with product engineers who are building the next generation of AI services. You will be their trusted advisor on reliability, helping shape their architecture and ensuring their products are operationally sound long before they hit production.

O pozici / o projektu

Původní popisek. 

We are looking for a seasoned Site Reliability Engineer to join the team responsible for the backbone of our global AI/ML services. This isn't your typical SRE role. You won't just be maintaining systems; you'll be the guardian of a massive, distributed AI compute platform that processes workloads at an incredible scale. You will ensure that our AI models and GPU-powered infrastructure are not just fast, but fundamentally reliable, observable, and built to last.

If you are passionate about building and operating large-scale systems and are excited by the unique challenges of the AI/ML world, this is the role for you.

Detail pracovní nabídky

  • Online nábor
  • Ihned
  • Práce plně na dálku

Výhody

  • Soukromá zdravotní péče
  • Sportovní balíček
  • Foreign languages classes
  • Life Insurance
  • Cafeteria system
Nabídka zveřejněnа Před 1 dnem