Sunday, 06 September 2026
Advertisement Advertise Your advert could be here Reach thousands of learners and ICT professionals across Rwanda. Contact us
Advertisement Opportunity Jobs, scholarships & hackathons Fresh openings from Rwandan job boards are pulled in every hour. See openings

Expert DevOps: scale, reliability and security

Run systems other people depend on — orchestration, observability, incidents, security and cost.

About this track

The top of the DevOps path, for someone who can already containerise an application and deploy it through a pipeline. This track is about what happens after that: orchestrating many services, being able to answer questions about a running system instead of guessing, deciding how reliable is reliable enough and alerting on that, handling an incident calmly and learning from it properly, securing the supply chain, and understanding what your infrastructure actually costs. It is deliberately opinionated about when not to reach for the complicated answer. Work through the lessons in order, then take the exam to earn your certificate.

Every lesson is written and hosted here on Yanjye — you never leave the site. Work through them in order, then sit the exam to earn your certificate.

Create a free account or log in to track your progress and earn the certificate.

Lessons

  1. 1
    What Kubernetes actually does

    The problem orchestration solves, and the price of solving it that way.

    ~30 min
  2. 2
    Pods, deployments, services and ingress

    The four objects that cover most of what you will actually write.

    ~35 min
  3. 3
    Configuration, secrets and stateless design

    The application properties that make scaling and replacement possible.

    ~30 min
  4. 4
    Scaling, requests and limits

    How capacity follows demand, and why the wrong limit causes the outage.

    ~30 min account needed
  5. 5
    Observability: logs, metrics and traces

    The difference between having data and being able to answer a question.

    ~35 min account needed
  6. 6
    SLOs, error budgets and alerts that mean something

    Deciding how reliable is reliable enough, and only waking people for that.

    ~30 min account needed
  7. 7
    Running an incident

    Roles, communication and the discipline of restoring service before understanding it.

    ~30 min account needed
  8. 8
    Postmortems that change something

    Why blame produces worse systems, and what to write instead.

    ~25 min account needed
  9. 9
    Backups and disaster recovery you have tested

    The one area where being wrong is unrecoverable.

    ~30 min account needed
  10. 10
    Security: supply chain and least privilege

    The attacks that actually happen, and the controls that actually stop them.

    ~35 min account needed
  11. 11
    What it costs, and why

    Reading a cloud bill and finding the few lines that are most of it.

    ~25 min account needed
  12. 12
    Platform engineering, and where to go from here

    Turning what you know into something a whole organisation can use.

    ~25 min account needed
Advertisement Yanjye Learn a new digital skill this week ICT, programming and professional courses with graded weekly assignments. Start free