We are looking for an experienced Cloud Infrastructure Manager to oversee and manage our enterprise IT infrastructure across Azure, Google Cloud Platform (GCP), and on-premises environments. The ideal candidate will ensure high availability, security, scalability, cost optimization, and operational excellence while leading infrastructure initiatives and supporting business-critical applications.
Key Responsibilities:
* Administer, maintain, and optimize Ubuntu/Linux servers across Azure, GCP, and on-premises environments.
* Manage cloud infrastructure including Azure Virtual Machines, Google Compute Engine, networking, storage, IAM, RBAC, VPNs, firewalls, and hybrid connectivity.
* Automate infrastructure provisioning and configuration using Terraform, Ansible, Bash, Python, Azure CLI, and Google Cloud CLI.
* Implement infrastructure security through patch management, system hardening, vulnerability remediation, access controls, and secrets management.
* Maintain monitoring, logging, alerting, backup, disaster recovery, and business continuity solutions using industry-standard tools.
* Support CI/CD pipelines, Docker-based deployments, and collaborate with DevOps, AI/ML, GIS, Data Engineering, and Application teams.
* Perform capacity planning, cloud cost optimization, resource rightsizing, and infrastructure performance tuning.
* Lead incident management, root cause analysis (RCA), production support, and infrastructure change management.
* Maintain infrastructure documentation, SOPs, architecture diagrams, inventories, and operational standards.
* Lead infrastructure engineers, coordinate with vendors, cloud providers, and ensure compliance with security and operational policies.
Preferred Skills :
* 4+ years of Strong hands-on experience with Ubuntu/Linux administration in production environments.
* Expertise in Microsoft Azure and Google Cloud Platform (GCP).
* Strong knowledge of networking (TCP/IP, DNS, VPN, NAT, Load Balancers, VPC, VNet, Firewalls).
* Hands-on experience with Terraform, Ansible, Infrastructure as Code (IaC), and automation.
* Experience with Linux security, IAM, RBAC, SSH, secrets management, and system hardening.
* Experience with monitoring and observability tools such as Azure Monitor, Google Cloud Monitoring, Prometheus, Grafana, ELK/OpenSearch, and Alert manager.
* Working knowledge of Docker and CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI.
* Experience with backup, disaster recovery, business continuity planning, and cloud cost optimization.
* Exposure to PostgreSQL, MySQL, Redis, Kafka, Cassandra, or similar infrastructure technologies is an added advantage.
* Excellent troubleshooting, documentation, leadership, vendor management, and stakeholder communication skills.