TRCI-26-05300
Site Reliability Engineer (SRE) – Observability/Monitoring Platform (Production Support)
Contract SRE/support role in Group Platform Services & Engineering, focused on administering and supporting a production observability/monitoring platform (Grafana LGTM and related tooling). You will act as custodian of the production environment, handle incidents/changes/releases/on-call, and engineer reliability through monitoring best practices, alert triage, RCA, and platform improvements. The role partners closely with engineers and architects to shape monitoring strategy, supports adoption across the group, and operates in a global agile team.
Position summary
- Location
- United Kingdom
- Workplace
- Hybrid
- Employment
- Contract
- Experience
- Minimum 4 years and Maximum 6 years
Role overview
Why This Role Matters.
Contract SRE/support role in Group Platform Services & Engineering, focused on administering and supporting a production observability/monitoring platform (Grafana LGTM and related tooling). You will act as custodian of the production environment, handle incidents/changes/releases/on-call, and engineer reliability through monitoring best practices, alert triage, RCA, and platform improvements. The role partners closely with engineers and architects to shape monitoring strategy, supports adoption across the group, and operates in a global agile team.
Your Impact
Deliver Enterprise Value
Help organisations solve complex business problems through modern technology, consulting expertise and measurable outcomes.
Collaboration
Work Across Teams
Collaborate with consultants, architects, engineers and client stakeholders throughout the project lifecycle.
Growth
Learn Continuously
Gain exposure to enterprise technologies, certifications, mentoring and real-world project experience.
Career Path
Grow With Ubique
Build a long-term consulting career with opportunities to take on greater responsibility and leadership over time.
Responsibilities
What You'll Be Doing.
Every role at Ubique contributes directly to solving meaningful business challenges for our clients.
Gain understanding of the various tools and frameworks that together provide observability and notification service to the organization and assist development and production support teams with queries / issues related to their usage of our platform.
Act as custodian of production environment and engage within the team and outside, if need be, towards building and maintaining robust, scalable, highly available production systems in accordance with our service level objectives
Preventing production incidents but when they do occur, performing effective incident and problem management and RCA to minimize downtime as well as possibility of recurrence.
Pushing out changes and releases to production environment reliably via effective change and release management
Quick and effective response to alerts before they become incidents, with an approach to prevent them from occurring ever again
Effectively triaging alerts, requests, emails such that things that needs attention get addressed first and in a timely manner in the order of their priority, the drivers for which should be production stability and user satisfaction.
Continuous and effective engagement with users, with the required empathy, providing the right guidance so as to provide a good customer experience
Collaborate in a global agile team environment using established support practices, participating in sprint planning, reviews, and continuous improvement initiatives
Build and maintain scalable, reliable monitoring solutions that support global infrastructure
Engage with engineers, architect towards contributing to architectural decisions that influence the future direction of observability platform
Champion observability best practices across the organization, helping teams leverage data-driven insights to improve system reliability and performance
Technology stack
Tools & Technologies.
The platforms and technologies you'll use to build modern, enterprise-grade solutions.
Grafana
LGTM
Loki
Grafana Tempo
Grafana Mimir
Prometheus
OpenTelemetry
Telemetry
Observability
Monitoring
Alerting
Linux
Python
Ansible
CI/CD
GitLab
Jenkins
Nexus
Kubernetes
EKS
Requirements
Skills & Experience.
We value curiosity, collaboration and continuous learning. If you don't meet every requirement but believe you can make an impact, we'd still love to hear from you.
Essential
Required Qualifications
Grafana administration or modern observability tool administration (2+ years)
Monitoring/observability platform support (enterprise scale)
Linux OS fundamentals and troubleshooting (2+ years)
Production support operations: incident management, problem management, change management, release management, on-call, alert response, request handling
Root cause analysis (RCA) and incident prevention mindset
Scripting/automation exposure: Python and/or Ansible
Basic cloud platform understanding
Basic CI/CD tool understanding (e.g., GitLab, Jenkins, Nexus, Ansible)
Strong analytical troubleshooting and mature judgment
Communication and interpersonal skills
Team collaboration (global/agile context)
Preferred
Nice to Have
OpenTelemetry standards understanding
Containerization: Kubernetes, EKS, Docker
ITIL knowledge
AI assistant tools usage with guardrails (e.g., Claude, Copilot)
RDBMS fundamentals and SQL (Sybase/MySQL/MSSQL)
Jira/Confluence
Familiarity with middleware and infra components (ActiveMQ, Solace, EMS, Tibco), web servers, load balancers, directory services
Experience supporting medium/large-scale production environments
Experience working with globally dispersed teams
What you'll gain
More Than Just A Job.
We're committed to helping every team member grow professionally, personally and technically while working on meaningful projects.
Global Exposure
Collaborate with international clients and multicultural teams on enterprise programmes.
Continuous Learning
Expand your expertise through mentoring, certifications and hands-on project experience.
Career Growth
Take ownership, develop leadership skills and grow your consulting career over time.
Flexible Working
Hybrid and remote collaboration designed around trust and delivering exceptional outcomes.
People First
Join a supportive culture where collaboration, respect and long-term relationships come first.
Enterprise Projects
Work on meaningful technology initiatives for leading organisations across industries.
Apply
Apply for Site Reliability Engineer (SRE) – Observability/Monitoring Platform (Production Support)
One page, about two minutes. We only ask for what we actually need to have a first conversation.