Search

Information Technology_USA - USA_Engineer

PublishedPublished: 4/18/2026
**Please strictly adhere to the following resume naming convention:
ALL CAPS, NO SPACES B/T UNDERSCORES

PTN_US_GBAMSREQID_CandidateBeelineID
i.e. PTN_US_9999999_SKIPJOHNSON0413

: /hr
MSP Owner: Subhashree samal
Location: - Phoenix,AZ
Duration: 6 months
skill id: 10908462
Job Title : Technical Manager/SRE Senior Lead

ROLE_DESCRIPTION
"Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.
(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
(Core) Prepare and maintain operational reports and documentation, including bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, and trend analysis.
(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues, ensuring effective communication and stakeholder alignment throughout the incident lifecycle.
Participate in Incident Management bridge calls, driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.
Monitor platform health and identify opportunities to improve system reliability, availability, observability, and operational efficiency.
Drive continuous service improvement initiatives by analyzing recurring incidents, identifying root causes, and recommending preventive actions.
Plan, coordinate, and facilitate Disaster Recovery (DR) exercises, ensuring readiness, documentation, and post-exercise review of outcomes."

SKILLS_REQUIRED
"Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.
(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
(Core) Prepare and maintain operational reports and documentation, including bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, and trend analysis.
(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues, ensuring effective communication and stakeholder alignment throughout the incident lifecycle.
Participate in Incident Management bridge calls, driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.
Monitor platform health and identify opportunities to improve system reliability, availability, observability, and operational efficiency.
Drive continuous service improvement initiatives by analyzing recurring incidents, identifying root causes, and recommending preventive actions.
Plan, coordinate, and facilitate Disaster Recovery (DR) exercises, ensuring readiness, documentation, and post-exercise review of outcomes."

ESSENTIAL_SKILLS
"Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.
(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
(Core) Prepare and maintain operational reports and documentation, including bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, and trend analysis.
(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues, ensuring effective communication and stakeholder alignment throughout the incident lifecycle.
Participate in Incident Management bridge calls, driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.
Monitor platform health and identify opportunities to improve system reliability, availability, observability, and operational efficiency.
Drive continuous service improvement initiatives by analyzing recurring incidents, identifying root causes, and recommending preventive actions.
Plan, coordinate, and facilitate Disaster Recovery (DR) exercises, ensuring readiness, documentation, and post-exercise review of outcomes."

KEYWORDS
Technical Manager/SRE Senior Lead

EXPERIENCE_RANGE_IN_REQUIRED_SKILLS : 10+ Years

Role Descriptions: Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.(Core) Prepare and maintain operational reports and documentation| including bridge updates| RCA tracking| incident trends| service availability| platform health metrics| monthly operational deliverables| and trend analysis.(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues| ensuring effective communication and stakeholder alignment throughout the incident lifecycle.Participate in Incident Management bridge calls| driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.Monitor platform health and identify opportunities to improve system reliability| availability| observability| and operational efficiency.Drive continuous service improvement initiatives by analyzing recurring incidents| identifying root causes| and recommending preventive actions.Plan| coordinate| and facilitate Disaster Recovery (DR) exercises| ensuring readiness| documentation| and post-exercise review of outcomes.
Essential Skills: Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.(Core) Prepare and maintain operational reports and documentation| including bridge updates| RCA tracking| incident trends| service availability| platform health metrics| monthly operational deliverables| and trend analysis.(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues| ensuring effective communication and stakeholder alignment throughout the incident lifecycle.Participate in Incident Management bridge calls| driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.Monitor platform health and identify opportunities to improve system reliability| availability| observability| and operational efficiency.Drive continuous service improvement initiatives by analyzing recurring incidents| identifying root causes| and recommending preventive actions.Plan| coordinate| and facilitate Disaster Recovery (DR) exercises| ensuring readiness| documentation| and post-exercise review of outcomes.
Desirable Skills:
Keyword:
Skills: Digital : Site Reliability Engineering (SRE)
Experience Required: