Astreya Logo

Astreya

Incident Response Analyst II

Posted 8 Days Ago
Be an Early Applicant
In-Office
Singapore, SGP
Entry level
In-Office
Singapore, SGP
Entry level
Monitor and investigate infrastructure, cloud, security, environmental, and threat alerts throughout the incident lifecycle. Triage and escalate incidents, coordinate resolver teams, serve as incident commander, document actions, conduct root cause analysis, and support emergency responses. The role monitors AWS, GCP, and Azure environments, identifies cloud anomalies and unauthorized access attempts, coordinates with engineering and security teams, and improves runbooks, playbooks, and operational processes.
The summary above was generated by AI

Responsibilities 

Incident & Problem Management

Analysts are responsible for the full lifecycle of incident management, from detection through to resolution and root cause analysis (RCA). This includes acting as incident commanders, maintaining SLAs, documenting actions, and providing insights to support continuous improvement efforts across teams and systems.

  • Investigate, report, and respond to alerts, incident response (war room, remote bridges).

  • Respond to incidents and critical situations in a calm, problem-solving manner, and conduct in-depth investigation of alerts.

  • Be the first line of defense using monitoring and automation tools to conduct investigation, classification, and triage, all within prescribed SLAs.

  • Provide deep understanding and intelligence of incident criticality and impact to resolver groups.

  • Ensure detailed records of alarm handling activities, including actions taken and resolutions in ticketing tools; file incident reports.

  • Act as incident commander during major incidents.

  • Understand internal/external communication methods and stakeholder responsibilities.

  • Support program managers and facilitate project deliverables, improving operational and engineering initiatives.

  • Conduct root cause analysis (RCA) to determine recurring problems.

  • Use in-depth questioning and analysis to determine the underlying cause of incidents or problems (Who, What, Where, When, Why).

  • Perform duties in compliance with SOPs, MOPs, Runbooks, and Playbooks.

Server, DCIM, Network and Traffic Alarms Operations

This function involves real-time monitoring of infrastructure alarms, determining the severity of alerts, escalating appropriately, and maintaining clear communications with resolver teams. It ensures uptime and system integrity across servers, network infrastructure, and environmental systems.

  • Continuously monitor alarm dashboards and systems.

  • Investigate and respond to alarms related to Network, Data Center Environment, Server Health, Facility Security, and Safety.

  • Identify and acknowledge incidents associated with alarms.

  • Assess incidents to determine their criticality and operational impact.

  • Engage resolver groups and escalate to higher tiers or management following established paths.

  • Maintain communication with teams, stakeholders, and incident responders.

  • Follow documented procedures to resolve incidents promptly and effectively.

  • Ensure accurate records of alarm handling and resolution activities in ticketing tools.

  • Comply with SOPs, MOPs, Runbooks, and Playbooks.

Threat Intelligence, Critical Event Management

Analysts monitor global threat feeds and operational alerts to protect ByteDance personnel and assets. Responsibilities include triaging alerts related to weather, security, travel, and regional instability, then coordinating appropriate response actions, escalating to law enforcement if necessary, and compiling response reports.

  • Monitor Everbridge Visual Command Center (VCC), InternationalSOS emails, and open-source tools for real-time incidents affecting ByteDance assets and travelers.

  • Monitor tools or queries for specific stakeholder requests.

  • Report on violence, severe weather, or threats to life, property, and assets.

  • Coordinate emergency responses, including with law enforcement if required.

  • Verify incident information accuracy through secondary sources.

  • Generate heatmaps to highlight affected areas during significant events.

  • Collaborate with security and operational teams for a coordinated response.

  • Implement incident containment and mitigation strategies.

  • Document incident details, response actions, and lessons learned.

  • Follow SOPs, MOPs, Runbooks, and Playbooks.

Cloud Incident Response and Monitoring

As hybrid environments become more critical to business operations, IRC Analysts will be expected to monitor and support both on-premises infrastructure and cloud-based systems. Analysts will assist in identifying and responding to cloud-related incidents across platforms such as AWS, GCP, and Azure. Responsibilities include:

  • Real-time monitoring of cloud infrastructure using tools such as AWS CloudWatch, Azure Monitor, and GCP Stackdriver.

  • Incident triage and escalation of alerts related to cloud-based services and resources (e.g., compute, storage, networking).

  • Coordination with Cloud Engineers and DevOps teams during cross-environment incidents to ensure rapid resolution and clear communications.

  • Identification and classification of cloud service anomalies, including misconfigurations, degraded services, and unauthorized access attempts.

  • Understanding of cloud-native architectures such as virtual private clouds (VPC), IAM, container orchestration (e.g., Kubernetes), and serverless functions.

  • Documentation of root cause analysis (RCA) and corrective actions for cloud incidents, feeding back into playbooks and runbooks.

  • Basic scripting and automation skills (Python, Bash, or PowerShell) for incident analysis and tooling.

  • Awareness of cloud security protocols, including encryption, IAM policies, and compliance standards like ISO 27001 and SOC 2.
     

Similar Jobs

15 Days Ago
In-Office
Singapore, SGP
Mid level
Mid level
Information Technology
Monitor alarms and handle cloud tickets; triage, escalate, and resolve incidents per SOPs and runbooks. Maintain shift continuity, documentation, and service reporting. Collaborate with SREs, vendors, and stakeholders, drive process improvements, and support major incident coordination and coaching for junior staff.
11 Days Ago
In-Office
Singapore, SGP
Junior
Junior
Information Technology
Operate 24/7 incident and alarm monitoring across servers, networks, DC infrastructure, and physical security. Triage and investigate alerts, act as incident commander during major events, perform RCA, coordinate threat and emergency responses (Everbridge/InternationalSOS), maintain ticketing records, and follow SOPs and playbooks to ensure uptime and safety.
Top Skills: Access Control SystemsAvigilonCctvDcimDnsEverbridge Visual Command Center (Vcc)GenetecGrafanaInternationalsosIp NetworksLenelLoad BalancingServer Health MonitoringTicketing Systems
17 Days Ago
In-Office
Singapore, SGP
Junior
Junior
Information Technology
Monitor and respond to facility and data center incidents 24x7: detect, triage, coordinate incident response, maintain ticketing records, run preliminary RCAs, manage communications with on-site teams and vendors, and support continuous improvement and reporting.
Top Skills: Automatic Transfer Switch (Ats)AvigilonBuilding Management Systems (Bms)CctvChilled Water SystemsCracCrahData Center Infrastructure Management (Dcim)Dcim/Epms Monitoring PlatformsElectrical Power Monitoring Systems (Epms)Environmental SensorsFire Alarm And Suppression SystemsGeneratorsGenetecHvacIncident Management SystemsJohnson ControlsLeak Detection SystemsLenelPower Distribution Units (Pdus)Schneider Electric EcostruxureSiemensTicketing SystemsUpsVertiv

What you need to know about the Singapore Tech Scene

The digital revolution has driven a constant demand for tech professionals across industries like software development, data analytics and cybersecurity. In Singapore, one of the largest cities in Southeast Asia, the demand for tech talent is so high that the government continues to invest millions into programs designed to develop a talent pipeline directly from universities while also scaling efforts in pre-employment training and mid-career upskilling to expand and elevate its workforce.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account