Ubique Systems

TRCI-26-04586

Site Reliability Engineer (Kubernetes/Cloud) - Contract

Contract SRE/DevOps engineer needed in Glasgow (or aligned) to run and optimize production Kubernetes clusters on a major public cloud (AKS/EKS/GKE). The role focuses on infrastructure as code with Terraform, CI/CD automation with Jenkins, strong incident response and root-cause analysis, and improving reliability through SRE practices (error budgets, capacity planning, blameless postmortems). You’ll also automate ops tasks with Python, enhance observability using Prometheus/Grafana/OpenTelemetry, and support GitOps workflows (e.g., Flux) while participating in on-call.

Position summary

Location
United Kingdom
Workplace
On-site
Employment
Contract
Experience
Minimum 4 years and Maximum 8 years
Apply now

Role overview

Why This Role Matters.

Contract SRE/DevOps engineer needed in Glasgow (or aligned) to run and optimize production Kubernetes clusters on a major public cloud (AKS/EKS/GKE). The role focuses on infrastructure as code with Terraform, CI/CD automation with Jenkins, strong incident response and root-cause analysis, and improving reliability through SRE practices (error budgets, capacity planning, blameless postmortems). You’ll also automate ops tasks with Python, enhance observability using Prometheus/Grafana/OpenTelemetry, and support GitOps workflows (e.g., Flux) while participating in on-call.

Your Impact

Deliver Enterprise Value

Help organisations solve complex business problems through modern technology, consulting expertise and measurable outcomes.

Collaboration

Work Across Teams

Collaborate with consultants, architects, engineers and client stakeholders throughout the project lifecycle.

Growth

Learn Continuously

Gain exposure to enterprise technologies, certifications, mentoring and real-world project experience.

Career Path

Grow With Ubique

Build a long-term consulting career with opportunities to take on greater responsibility and leadership over time.

Responsibilities

What You'll Be Doing.

Every role at Ubique contributes directly to solving meaningful business challenges for our clients.

01

Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Azure AKS, AWS EKS, or Google GKE).

02

Implement and maintain infrastructure as code using tools such as Terraform.

03

Collaborate with development and operations teams to improve system reliability and deployment automation.

04

Build and maintain CI/CD pipelines using Jenkins or similar tools.

05

Troubleshoot production issues, conduct root cause analysis, and implement preventive measures.

06

Automate operational tasks using Python or other scripting languages.

07

Contribute to observability and monitoring improvements using modern tools and best practices.

08

Participate in on-call rotations and incident response processes.

09

4–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles.

10

Strong hands-on experience managing Kubernetes clusters in production (AKS/EKS/GKE).

11

Proficiency with Terraform and cloud infrastructure automation.

12

Practical experience with Jenkins and CI/CD pipeline management.

Technology stack

Tools & Technologies.

The platforms and technologies you'll use to build modern, enterprise-grade solutions.

01

Kubernetes

02

AKS

03

EKS

04

GKE

05

Azure

06

AWS

07

GCP

08

Terraform

09

Infrastructure as Code

10

IaC

11

Jenkins

12

CI/CD

13

SRE

14

Site Reliability Engineering

15

Incident management

16

Blameless postmortems

17

Error budgets

18

Capacity planning

19

Python

20

Prometheus

Requirements

Skills & Experience.

We value curiosity, collaboration and continuous learning. If you don't meet every requirement but believe you can make an impact, we'd still love to hear from you.

Essential

Required Qualifications

4–8 years in SRE/DevOps/Cloud Infrastructure roles

Production Kubernetes operations (AKS or EKS or GKE)

Public cloud experience (Azure or AWS or GCP)

Terraform (infrastructure as code)

Jenkins and CI/CD pipeline management

SRE principles (incident management, blameless postmortems, capacity planning, error budgets)

Troubleshooting, root cause analysis, and preventive remediation

Scripting/programming (Python preferred or similar)

Observability/monitoring with Prometheus and/or Grafana and/or OpenTelemetry

On-call and incident response participation

Strong written and verbal communication

Preferred

Nice to Have

GitOps practices

Flux (GitOps tool)

Broader CI/CD tooling beyond Jenkins (e.g., similar tools)

Additional scripting languages beyond Python

What you'll gain

More Than Just A Job.

We're committed to helping every team member grow professionally, personally and technically while working on meaningful projects.

Global Exposure

Collaborate with international clients and multicultural teams on enterprise programmes.

Continuous Learning

Expand your expertise through mentoring, certifications and hands-on project experience.

Career Growth

Take ownership, develop leadership skills and grow your consulting career over time.

Flexible Working

Hybrid and remote collaboration designed around trust and delivering exceptional outcomes.

People First

Join a supportive culture where collaboration, respect and long-term relationships come first.

Enterprise Projects

Work on meaningful technology initiatives for leading organisations across industries.

Apply

Apply for Site Reliability Engineer (Kubernetes/Cloud) - Contract

One page, about two minutes. We only ask for what we actually need to have a first conversation.

Your CV

Optional. A line or two is plenty.