### The Apps Support Intmd Analyst is a developing professional role. Deals with most problems independently and has some latitude to solve complex problems. Integrates in-depth specialty area knowledge with a solid understanding of industry standards and practices. Good understanding of how the team and area integrate with others in accomplishing the objectives of the subfunction/ job family. Applies analytical thinking and knowledge of data analysis tools and methodologies. Requires attention to detail when making judgments and recommendations based on the analysis of factual information. Typically deals with variable issues with potentially broader business impact. Applies professional judgment when interpreting data and results. Breaks down information in a systematic and communicable manner. Developed communication and diplomacy skills are required in order to exchange potentially complex/sensitive information. Moderate but direct impact through close contact with the businesses' core activities. Quality and timeliness of service provided will affect the effectiveness of own team and other closely related teams.
Responsibilities:
###
### System Stability and Availability: Ensure the reliability, availability, and scalability of our Java-based applications and databases, running on both on-premise VMs and cloud environments.
Infrastructure Management: Manage and apply Infrastructure as Code (IaC) principles using tools like Ansible to maintain consistent and repeatable environments.
CI/CD and Automation: Oversee and enhance existing CI/CD pipelines to support reliable software delivery. Identify opportunities for automation to reduce toil and improve operational efficiency.
Monitoring and Observability: Manage and enhance comprehensive monitoring, logging, and alerting solutions (e.g., Open Telemetry, Grafana, ELK stack, GCO) for proactive identification, troubleshooting, and resolution of issues.
Incident Management: Lead incident response, conduct blameless post-mortems, and drive the implementation of corrective actions to prevent recurrence and improve system resilience.
Database Reliability: Uphold the reliability and performance of our database systems (e.g., BigData, MySQL, Oracle). Focus on performance tuning, backup and recovery strategies, and disaster recovery preparedness.
Performance Analysis: Proactively identify, analyze, and address performance bottlenecks across the application and infrastructure stack to ensure optimal performance.
Security and Compliance: Apply and enforce security best practices throughout the infrastructure and application lifecycle. Ensure compliance with industry standards and internal policies.
###
### Leveraging AI in Day-to-Day Tasks:
###
### AI-Powered Monitoring: Utilize AI and machine learning models to enhance monitoring capabilities, enabling predictive alerting and anomaly detection to identify potential issues before they impact users.
Automated Root Cause Analysis: Leverage AI-driven tools to accelerate root cause analysis during incidents by analyzing logs, metrics, and traces to pinpoint the source of the problem.
Intelligent Automation: Utilize AI-powered automation to handle complex operational tasks, such as intelligent resource scaling, automated remediation of common issues, and predictive capacity planning.
ChatOps and AI Assistants: Work with AI-powered chatbots and assistants within operational workflows to streamline communication, handle routine queries, and gain real-time insights during incident response.
###
### Collaboration and Leadership:
###
### Cross-Functional Collaboration: Work closely with development, QA, and security teams to foster a culture of reliability and ensure services meet high standards for availability and performance.
Mentorship: Mentor junior SREs and share your expertise to help grow the team's capabilities in system stability and monitoring.
Technical Guidance: Provide expert technical guidance on...
Want jobs like this matched to you?
Swoopd scores fresh postings against your résumé so you only see the matches that matter.