Server Administrator
Date: 6 Aug 2026
Location: IN
Company: Digital Edge DC
Who we are:
Where performance meets sustainability, Digital Edge is a leading data center platform in Asia Pacific, delivering sustainable, next-generation, AI-ready digital infrastructure at scale. Headquartered in Singapore and backed by Stonepeak, we operate a rapidly expanding platform with over 30 data centers and 1.8GW of secured IT power across key markets.
Our mission is to build the foundation for the world’s digital future, helping organizations to grow sustainably and empowering the populations they serve. We differentiate through execution – delivering at speed and at scale while combining global standards with deep local expertise. This enables us to provide consistent, sustainable and high-performance infrastructure for hyperscalers and enterprises across the region.
As Asia Pacific’s digital transformation accelerates, Digital Edge is building the infrastructure that powers the next wave of AI and cloud growth. Joining Digital Edge means being part of a high-growth platform at the forefront of AI and cloud infrastructure, where you can make a tangible impact on Asia-Pacific’s digital future.
Our values:
- Respect: We embrace diversity and collaboration.
- Innovation: We share ideas and solve problems.
- Strive: We are driven and determined.
- Excellence: We seek to deliver the best.
- Responsibility: We do what’s right for our people and the planet.
What we need:
Reporting to the Critical IT Director, this role is the regional owner of Digital Edge’s mission-critical on-premises server infrastructure across our data centre portfolio. This is an estate of more than 100 physical servers, supporting more than 180 critical virtual machines. It unites deep engineering ownership of high-performance server clusters built on the hyperconverged infrastructure (HCI) concept with 24/7 operational excellence. The role owns the physical server hardware and the hypervisor layer that underpin the HCI platform; the operating systems and applications running within the virtual machines are managed by their respective owners and sit outside the scope of this role. The platform is vendor agnostic, and the role demands a deep understanding of the underlying technology stack of compute, virtualisation, software-defined storage and the management planes that bind them. The emphasis is not on deploying the latest and greatest virtualization technology, but on consistency of deployment and a disciplined operational mindset across the portfolio.
Key responsibilities:
- Performance, Optimization & Capacity Engineering
- Own end-to-end performance engineering of high-performance server clusters, establishing baselines and tuning compute, virtualisation and storage layers to eliminate bottlenecks across the stack.
- Lead capacity planning and forecasting for the regional server estate, modelling workload growth and ensuring headroom for expansion without over-provisioning.
- Define and track KPIs for cluster health, utilisation and efficiency, driving continuous optimisation of resource allocation and workload placement.
- Benchmark and evaluate new server hardware generations and HCI software releases, validating price-performance before fleet-wide adoption.
- Infrastructure Strategy & Resilience Architecture
- Own the regional server infrastructure architecture and design standards, built on the hyperconverged infrastructure (HCI) concept while remaining vendor agnostic across hardware and software platforms.
- Design and continually enhance high-availability, clustering, replication, backup and disaster-recovery architectures to meet mission-critical availability targets.
- Develop the multi-year technology roadmap for the server platform, prioritizing consistent, standardized deployments and operational stability over adoption of the newest technology, with predictable lifecycle management and hardware refresh cycles.
- Act as the subject-matter expert (SME) and technical authority for server infrastructure across the regional portfolio, engaging early in data centre design and expansion projects to inform compute deployment decisions.
- 24/7 Operations, Availability & Security
- Run day-to-day operations and maintenance of production server clusters at the hardware and hypervisor layers within a 24/7 operational model, including provisioning, patching, configuration, change, asset and fault management.
- Deploy and operate supporting platform services alongside the core clusters, such as backup infrastructure (e.g. Proxmox Backup Server).
- Serve as the escalation point for major incidents affecting the server estate, providing 24/7 standby on rotation as the last-resort escalation tier, and leading root-cause analysis and driving permanent corrective actions.
- Own security hardening of the server estate, including hypervisor and firmware baselines, vulnerability remediation, access control and audit readiness.
- Ensure comprehensive monitoring, alerting and observability coverage across the fleet, enabling proactive detection of capacity, performance and availability risks.
- Automation, Tooling & Standardization
- Build and maintain infrastructure automation and infrastructure-as-code using tools such as Ansible, Terraform and Python to streamline provisioning, configuration and operations.
- Drive standardisation of server builds, golden images, cluster configurations and operational runbooks across the portfolio.
- Propose and implement system enhancements that improve the performance, reliability and efficiency of the server infrastructure.
- Vendor & Stakeholder Management
- Manage OEM, software and support vendors, including evaluation, coordination and escalation, while preserving the platform’s vendor-agnostic posture.
- Serve as the primary interface for the internal teams whose virtual machines run on the platform, managing onboarding, capacity requests and resource allocation against agreed service levels.
- Communicate proactively with virtual machine owners on planned maintenance windows, platform changes and incidents affecting their workloads, setting clear expectations on service impact and restoration.
- Define and document the shared-responsibility boundary between the platform and virtual machine owners, keeping hypervisor-level and guest-level duties clearly delineated.
Actively champion and implement the organization’s policies on health & safety, environment, energy, quality, information security, and business continuity, ensuring adherence to incident reporting, legal, and regulatory requirements.
The successful candidate:
- Diploma/ Degree in computer science, engineering or equivalent experience
- 5+ years of experience in server and compute infrastructure engineering and operations within mission-critical environments (data centre, colocation, cloud or service provider)
- Deep hands-on expertise with HCI and virtualisation platforms (e.g. Proxmox VE with Ceph, VMware vSphere/vSAN, Nutanix, Hyper-V), including cluster design, high availability and live migration, with a vendor-agnostic mindset and familiarity with supporting services such as Proxmox Backup Server and Proxmox Mail Gateway
- Strong understanding of enterprise server hardware, including firmware lifecycle, out-of-band management (iDRAC, iLO, BMC) and facilities considerations such as power, cooling and rack density
- Experienced in administering Linux at the hypervisor and storage layers at scale, with a strong understanding of Ceph (RBD, CephFS, replication and erasure coding) and data protection (backup, replication, disaster recovery)
- Hands-on experience with automation and infrastructure-as-code (e.g. Ansible, Terraform, Python), and with monitoring, observability and capacity-management tooling
- Ability to participate in a 24/7 rotational standby as the last-resort escalation point, and to support planned maintenance windows
- Excellent analytical and problem-solving skills, with the ability to work effectively with clients, senior management, staff and vendors
Preferred (advantageous):
- Professional certifications, such as VMware VCP/VCAP, Nutanix NCP/NCM, Microsoft (Windows Server / Azure Local) or Red Hat (RHCE) certifications.
Join us to shape the future of Digital Infrastructure in Asia. Apply now for a confidential career discussion.