Back to board
relanto·7 days ago

Cloud Infrastructure Operations Specialist

Bengaluru, IndiaFull-timeMid · 2-5 years

About this role

Job Title: Cloud Infrastructure Operations Specialist
Role Overview
The Cloud Infrastructure Operations Specialist plays a critical role in managing, allocating, and monitoring cloud infrastructure resources across global data centers. This position requires a highly adaptable professional who can manage complex capacity requests, ensure seamless service operations, and maintain fleet health. The ideal candidate thrives in a collaborative environment, is comfortable working outside standard business hours to support global operations and possesses strong technical troubleshooting skills.
Key Responsibilities
• Resource Allocation & Triage:
Review, evaluate, and fulfils incoming capacity requests for large-scale internal services (such as distributed compute clusters, cloud storage systems, and globally managed databases) using enterprise ticketing and tracking systems.
• Policy Enforcement & Validation:
Verify existing compute and storage presence in requested geographic data center zones and enforce infrastructure policies by pushing back on unsupported infrastructure scaling requests.
• Emergency Incident Mitigation:
Handle urgent capacity shortages and production-impacting incidents by executing emergency quota allocations, resource loans, or scaling operations using infrastructure command-line tools.
• Fleet Optimization & Monitoring:
Monitor cluster health, process hardware consolidation and datacenter lifecycle requests, and approve automated infrastructure changes to optimize capacity and preserve overall fleet health.
• Cross-Functional Collaboration:
Participate in special ad-hoc capacity reclaim projects and ensure resources are properly tracked, audited, and returned to the shared resource pool when no longer needed.
Requirements & Qualifications
• Shift Flexibility:
Willingness to work in night shifts to provide continuous support for global infrastructure and operations.
• High Availability:
Highly reliable and available to respond to critical escalations or urgent capacity shortages as needed.
• Team Player:
Excellent teamwork and collaboration skills, with the ability to communicate effectively with engineering teams, resource managers, and cross-functional stakeholders.
• Technical Proficiency (SQL & Unix):
Hands-on experience with Unix/Linux environments and command-line interfaces (CLI) for executing system operations, along with a strong understanding of SQL for database querying, data analysis, and reporting.
• Analytical Skills:
Ability to analyze large datasets (e.g., using advanced spreadsheets) to calculate resource savings, monitor usage trends, and track inventory effectively.
• Problem-Solving:
Strong critical thinking skills with the ability to quickly evaluate technical requests and apply proper operational guidelines to ensure system stability.