Senior Software Engineer (vMetal)
name
40d ago
0DevUnited Stateshimalayas
Senior-Software-EngineerSenior-Software-Development-EngineerSenior-Lead-Software-EngineerSr.-Software-EngineerSenior-Staff-Software-EngineerLead-Senior-Software-EngineerSenior-Software-Engineer---C#Senior
Job Description
As vCluster Labs' Senior Software Engineer (vMetal), you aren't just running servers; you are turning a rack of bare metal into a programmable platform. In this role, you will build the engineering work behind vMetal, the systems that discover, provision, configure, and manage physical hardware so our customers can spin up tenant clusters on top of it. You will sit at the rare intersection of out-of-band server management and modern Go services, and you will be one of the first engineers to help us grow this team. You will report to the VP of Engineering.As a Senior Software Engineer (vMetal), your role will include:vMetal: You will contribute to architecture decisions, help drive the roadmap for bare metal provisioning and lifecycle management, and hold a high bar for code review and design quality on the team.Bare Metal Programmability: Build the Go services that turn raw hardware into APIs our customers can consume. You will design and ship the systems that drive Redfish, IPMI, and PXE workflows in production, not glue scripts, real services with clean interfaces and solid tests.Hardware Lifecycle Automation: Own how servers get discovered, inventoried, provisioned, configured, and reclaimed. You will eliminate manual intervention from the day-2 path and design for hardware that fails in surprising ways.Cross-Generational Architecture: Translate between traditional out-of-band server management and modern Kubernetes-native patterns. You will contribute to where the abstractions live and how vMetal exposes hardware to tenant clusters cleanly.Customer-Driven Reliability: Partner with customer engineering and the broader platform team to debug, harden, and ship against real production workloads. You will be on-call for the systems you build and you will treat reliability as a first-class deliverable.This role could be a fit for you if you bring:Go Fluency: You write production Go for a living. You can design clean services, APIs, and libraries, not just script around someone else's code.Bare Metal Operations Depth: You have shipped systems that drive servers via Redfish and IPMI in production, and you understand PXE boot end-to-end. You have debugged what happens when a BMC lies to you.Bare-Metal-as-a-Service Background: You have built or operated bare-metal-as-a-service offerings, an AI Cloud, or a hyperscaler bare metal team where infrastructure was the product.Bridging Generations: You hold both server management and modern API design in your head and can design between them. You are comfortable in IPMI, iDRAC, and ILO consoles, and equally comfortable shipping a Go controller.Operator's Instinct: You think about failure modes, telemetry, and recoverability before you ship. You have been on-call for the systems you have built and you treat that as a feature, not a tax.Bonus points for:Kubernetes Fluency: Comfort with Kubernetes internals, controllers, or operators, enough to design how bare metal hosts integrate cleanly with tenant clusters.AI Cloud / Hyperscaler Time: Production experience inside an AI Cloud, hyperscaler, or large platform team where bare metal scale was non-negotiable.Open Source Contributions: Meaningful contributions to projects in the bare metal, provisioning, or Kubernetes ecosystem, such as Tinkerbell, Metal3, Cluster API, or Ironic.About vCluster LabsWe are a venture-backed tech startup and the company pioneering Kubernetes virtualization for the AI era. We raised +$30M from top-tier VCs such as Khosla Ventures (first investor in OpenAI, GitLab, Stripe, Doordash) and are in a hyper-growth phase looking for motivated people to complement our team. Our headquarters are in San Francisco (Salesforce Tower), but our team is distributed around the globe and we have a remote-first work culture.We are the leading platform for operating GPU infrastructure, enabling AI Cloud providers to deliver a hyperscaler-like experience to their customers and AI factories that need to build that same experience for their internal teams. Our platform delivers the full operational stack operators need to run their GPU data centers — managed Kubernetes, fast isolated tenant provisioning, and automated node provisioning and lifecycle management — enabling them to accelerate time to value, reduce operational burden, and maximize the ROI of every GPU.We're the company behind vCluster, an open-source technology for virtualizing Kubernetes (10k+ GitHub stars, 40M+ virtual clusters created since 2021). Open source is part of our DNA. At KubeCon North America 2025, we launched our Infrastructure Tenancy Platform for AI — a Kubernetes-native framework purpose-built for running AI, ML, and GPU-intensive workloads anywhere, with an NVIDIA-validated reference architecture for DGX systems.BenefitsWe offer the following benefits:Competitive Salary: We offer a competitive compensation package, including equity.Platinum-Level Insurance: Health, dental, vision, and life Insurance, including plans for you and eligible de
