Skip to content
Skip to content
DevOps Jobs
NextEra Energy

Senior Observability Engineer (SRO)

NextEra Energy

Location
Onsite (Juno Beach, FL · Plantation, FL)
Employment
Full-time
Level
Senior Level
Posted 1 day ago

About the Role

Florida Power & Light, the largest electric utility in the U.S., is seeking a Senior Observability Engineer to advance SRE practices and enterprise observability. This role focuses on improving service health, reducing operational noise, and driving data-driven reliability across IT operations.

Skills

SRE principles Observability ServiceNow ITOM Splunk ScienceLogic AppDynamics SLIs SLOs Automation Incident management Metrics Logs Traces Synthetic monitoring Troubleshooting Data-driven operations

Full job details

Requisition ID:  96990 

Florida Power & Light Company is the largest electric utility in the U.S., providing reliable energy to nearly 12 million Floridians. With one of the nation’s most fuel-efficient, cost-effective power generation fleets and industry-leading reliability, we’re redefining what’s possible in energy. Want to be part of something powerful? Join our outstanding team and help shape the future of energy.

 

Position Specific Description

The Senior Observability Engineer / Architect will help advance modern SRE practices across enterprise IT operations, with a focus on service-level visibility, observability, actionable alerting, automation, runbook maturity, and proactive service-health management. This role will partner across Observability, Event Management, infrastructure, application, and ServiceNow teams to improve reliability engineering standards, strengthen monitoring coverage, reduce operational noise, and support the transition from reactive incident response to data-driven, service-health operations. 

The position will provide a technical connection point across Information Technology, infrastructure, application, cloud, ServiceNow, and operations teams to improve reliability, observability, and operational readiness. This role will support the development and execution of observability standards, service health practices, alerting improvements, automation opportunities, and reliability engineering patterns that help teams detect, understand, and resolve service issues more effectively.

Project Execution / Analytical Thinking / Problem Solving 

• Assist in designing, implementing, and operating enterprise observability capabilities across metrics, logs, traces, events, synthetic monitoring, dashboards, and service-health views. 
• Apply modern SRE principles, including SLIs, SLOs, error-budget thinking, toil reduction, automation, incident learning, and reliability-focused engineering practices. 
• Partner with infrastructure, application, cloud, database, network, storage, and operations teams to define monitoring requirements, alert thresholds, escalation paths, and service-health indicators. 
• Support observability platform capabilities across tools such as ScienceLogic, ServiceNow ITOM/Event Management, Splunk, cloud-native monitoring platforms, AppDynamics, synthetic monitoring, and related technologies. 
• Improve alert quality by helping ensure alerts are actionable, properly routed, associated with the correct configuration item or service, and supported by clear response guidance. 
• Assist in aligning operational events to ServiceNow Event Management, including event ingestion, alert correlation, suppression logic, incident creation criteria, and notification workflows. 
• Contribute to service-level visibility by supporting dashboards, scorecards, service maps, dependency views, ownership models, and operational health reporting. 
• Analyze recurring incidents, monitoring gaps, alert patterns, and operational trends to identify reliability improvement opportunities. 
• Develop and maintain runbooks, knowledge articles, technical documentation, monitoring standards, and operational handoff materials. 
• Support automation opportunities that reduce manual effort, improve triage consistency, and accelerate restoration while maintaining appropriate governance and controls. 
• Collaborate with teams during incidents and problem reviews to improve detection, escalation, root-cause analysis, and long-term prevention. 
• Help advance observability maturity through practical adoption of standards such as OpenTelemetry where appropriate, along with consistent telemetry collection and platform integration practices. 
• Respond to complex operational scenarios where standard procedures have not resolved the issue and provide technical analysis to support restoration and prevention. 
• Continuously evaluate observability practices, platform effectiveness, data quality, and service readiness to improve reliability outcomes across the enterprise. 


Skills / Preferred Qualifications 

• Strong understanding of SRE principles and reliability engineering practices. 
• Experience with metrics, logs, traces, events, and dashboards. 
• Knowledge of observability tools such as ScienceLogic, ServiceNow ITOM, Splunk, or AppDynamics. 
• Ability to define SLIs, SLOs, and service-health indicators. 
• Experience improving alert quality, routing, and correlation. 
• Strong troubleshooting and problem-solving skills. 
• Scripting or automation experience to reduce manual effort. 
• Ability to collaborate across infrastructure, application, cloud, and operations teams.