About Nscale
Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.
We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.
About the Role
We're hiring a Staff Software Engineer to build the software, automation, and control-plane capabilities that manage Nscale's fleet of AI infrastructure at scale. Your work will improve the acceptance, performance, and scalability of our AI and high-performance computing environments — driving higher availability, faster capacity delivery, and lower operational load as Nscale grows into one of the world's leading neo-cloud providers.
This is a senior individual-contributor role for an engineer who enjoys solving hard infrastructure problems at the intersection of software, GPUs, networking, and large-scale operations. You will have the autonomy to investigate problems, learn quickly, innovate, and deliver improvements wherever they create meaningful impact for the team and the platform.
You will work closely with teams across Nscale — including Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware engineering — to translate operational challenges into reliable, scalable software. You will not need to own every component to make a difference: strong engineers identify opportunities, build a compelling case for a solution, and work with the right partners to deliver it.
NOTE: We are hiring for various senior experience levels. The final leveling for the role will be based on your overall work experience, experience in AI Infra domain and interview feedback.
What You'll Be Doing
- Lead the architecture, roadmap, and implementation of workflow automation and fleet-management systems, balancing scalability, reliability, and maintainability.
- Build and operate production-grade software, services, APIs, and automation that manage the lifecycle of GPU compute and supporting network infrastructure.
- Own end-to-end workflows for fleet inventory, provisioning, configuration, hardware and firmware lifecycle management, validation, health monitoring, remediation, capacity, and reliability at scale.
- Investigate complex production issues across hardware, GPUs, operating systems, networks, schedulers, and services; turn findings into durable software improvements rather than recurring manual work.
- Build safe, observable, and auditable control-plane workflows that give operators clear visibility and reliable ways to act.
- Establish engineering standards for reliability, observability, testing, CI/CD, security, incident response, and operational readiness. Use SLOs, telemetry, alerting, and postmortems to drive continuous improvement.
- Partner with Deployment, AI Infrastructure Support, Data Centre Operations, Platform, SRE, Network, and hardware teams to translate operational needs into robust, scalable automation.
- Influence the evolution of adjacent systems and services through sound technical judgment, clear communication, and practical solutions.
- Assess the impact of new hardware programmes on the software stack and ensure fleet-management capabilities are ready to support them.
- Lead technical design reviews and incident deep-dives; mentor other engineers and raise the engineering bar across the organization.
- Use AI-assisted development tools to increase delivery leverage while maintaining a high bar for correctness, security, and operational safety.
About You
- 8+ years of experience building and operating large-scale infrastructure applications, platform services, cloud systems, or equivalent production systems.
- A Bachelor's degree in Computer Science, Computer Engineering, a relevant technical field, or equivalent practical experience.
- A strong software-engineering foundation in Python and/or Go, Java, C++, or similar languages, including API design, testing, code review, and production debugging.
- Deep understanding of Linux, distributed systems, networking fundamentals, and systems performance; you are comfortable working across stateful and stateless services.
- Experience designing and operating reliable automation or control-plane systems for complex infrastructure, large fleets, cloud platforms, or hardware lifecycle management.
- Proven ability to take ambiguous technical problems from architecture through implementation and production operation, while influencing peers and stakeholders without relying on formal authority.
- Hands-on experience with observability, monitoring, metrics, logs, tracing, alerting, incident response, capacity planning, and performance analysis.
- Strong communication skills and sound technical judgment. You can explain trade-offs clearly, build alignment, and move work forward in a fast-changing environment.
- A curious, pragmatic, high-ownership mindset. You enjoy finding the underlying cause of difficult problems and building the simplest durable solution.
Strong Candidates Will Have
- Direct experience with AI, GPU, HPC, or large-scale cloud infrastructure, including NVIDIA GPUs, CUDA, NVLink/NVSwitch, NCCL, and workload schedulers such as Slurm and Kubernetes.
- Experience with high-performance datacentre networking, including InfiniBand, RoCE, Ethernet fabrics, routing, congestion control, topology-aware systems, or GPU Direct RDMA.
- Experience with bare-metal lifecycle automation and infrastructure management tools such as Redfish, IPMI, PXE, MAAS, Ironic, NetBox, DCIM, OpenStack, or equivalent systems.
- Experience with workflow orchestration and reliable automation systems such as Temporal, Airflow, Prefect, or event-driven architectures.
- Experience with Kubernetes, containers, infrastructure as code (Terraform, Pulumi, Ansible), and public-cloud or private-cloud platforms.
- Experience with observability platforms and high-cardinality telemetry, such as Prometheus, Grafana, OpenTelemetry, ELK, or equivalent.
- Experience with hardware qualification, burn-in, validation, fleet health, or automated remediation for servers, GPUs, or network equipment.
- A track record of technical leadership: setting direction, defining reusable patterns, developing other engineers, and improving the effectiveness of multiple teams.
What We Can Offer You
At Nscale, you'll find a collaborative, supportive, and innovative environment where your contributions spark real impact. We're building something extraordinary, and we want you at the core.
- Highly competitive package (base + equity) with reviews every 12 months.
- Join the fastest-growing tech startup, your chance to push boundaries, collaborate with brilliant minds, and make your mark on cutting-edge AI.
- Expect a dynamic progression plan tailored to your ambitions. Grow by trying new things, leading, challenging the status quo, and owning your impact, always with our full support.
Equal Opportunities Statement
We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.
If there’s anything we can do to accommodate your specific situation, please let us know.
The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.