About this role
Job title: Subject Matter Expert (Support&Ops)
About the Role Subject Matter Expert (Support&Ops) to lead and implement enterprise observability using Grafana and related frameworks. Responsible for designing, deploying, and optimizing monitoring solutions and ensuring reliable observability across large-scale environments.
What You'll Do
- Administer Grafana including dashboard development, alert configuration, RBAC, user management, and data source integration.
- Install, configure, troubleshoot, upgrade Grafana plugins; optimize performance for enterprise-scale monitoring.
- Design and maintain observability solutions using Grafana Alloy, OpenTelemetry.
- Configure Grafana Alloy pipelines, telemetry collection/log forwarding, relabeling, filtering, and performance tuning.
- Manage BindPlane administration: collector deployment, gateway configuration, telemetry routing, load balancing, high availability, and troubleshooting.
- Ingest telemetry via pipelines from on-prem and cloud infrastructure into centralized observability platforms.
- Administer GCP services with hands-on GKE cluster administration, workload deployment, pod management, scaling, and troubleshooting.
- Use Google Cloud Monitoring tools (Metrics Explorer, Logs Explorer, dashboards, alert policies) and adopt observability best practices.
- Perform strong Kubernetes administration (deployments, services, ingress, daemonsets, statefulsets, namespaces, resource management, troubleshooting).
- Manage and monitor AKS environments; implement observability for containerized workloads.
- Apply knowledge of Azure cloud services, networking, identity management, and infrastructure monitoring.
- Use Ansible for infrastructure automation, configuration management, deployment automation, and operations tasks.
- Develop automation and monitoring solutions with Python and Shell Scripting.
- Integrate monitoring platforms with ServiceNow, REST APIs, webhook-based alerting, SQL, and third-party enterprise applications.
- Conduct root cause analysis, capacity planning, performance optimization, and reliability improvements for large-scale monitoring platforms.
- Support enterprise observability environments with thousands of monitored servers, applications, and cloud-native workloads.
- Communicate effectively with stakeholders; document findings and improvements.
What We're Looking For
- Strong hands-on Grafana administration: dashboards, alerting, RBAC, user management, data source integration.
- Expertise with Grafana Alloy, OpenTelemetry; plugin installation, configuration, troubleshooting, upgrades, and performance optimization.
- Experience designing/maintaining observability solutions with Grafana Alloy, Grafana, OpenTelemetry.
- Hands-on Grafana Alloy configuration, telemetry pipelines, log/metric forwarding, relabeling, filtering, performance tuning.
- BindPlane administration: collector deployment, gateway configuration, telemetry routing, load balancing, high availability, troubleshooting.
- Experience ingesting telemetry from on-prem/cloud infrastructure into centralized observability platforms.
- GCP services with GKE cluster administration, workload deployment, pod management, scaling, troubleshooting.
- Experience with Google Cloud Monitoring tools (Metrics Explorer, Logs Explorer, dashboards, alerting policies).
- Strong Kubernetes administration: deployments, services, ingress, daemonsets, statefulsets, namespaces, resource management, troubleshooting.
- Experience managing AKS environments and observability for containerized workloads.
- Knowledge of Azure cloud services, networking, identity management, and infrastructure monitoring.
- Ansible for automation; Python and Shell scripting for monitoring and integrations.
- REST APIs, ServiceNow, webhook-based alerting, SQL, and third-party enterprise apps integration.
- Linux system administration, troubleshooting, process management, networking fundamentals, performance analysis.
- Ability to perform root cause analysis, capacity planning, performance optimization, reliability improvements for large-scale monitoring.
- Experience supporting enterprise observability environments with thousands of servers and cloud-native workloads.
- Excellent analytical, troubleshooting, documentation, and stakeholder communication skills.
Nice to Have
- Prometheus familiarity; experience with Grafana Alloy configuration and telemetry pipelines is a plus.
Compensation & Benefits
- Not disclosed in posting.