Senior System Software Engineer

LB119
  • $250,000
  • Mountain View, CA
  • Permanent

System Software Engineer, Node & Cluster Management

Hybrid | Mountain View, CA | Tuesday to Thursday on-site

$250k+ base + RSUs


I’m working with a well-funded AI semiconductor startup building a new processor architecture for frontier-scale language models.

The host software team owns the stack that makes its custom AI hardware usable, from low-level Linux interfaces through node and cluster management. This role will help define how a new compute platform is monitored, controlled, updated, and recovered at datacenter scale.


You’ll work on problems such as:

  • Building node-level management services for health, telemetry, inventory, and control
  • Designing cluster-management and failover capabilities that minimize downtime
  • Developing REST APIs and CLI tools for diagnostics, firmware updates, and device recovery
  • Creating unified management across host software and OpenBMC firmware
  • Extending management capabilities from individual nodes to racks and clusters
  • Debugging issues across APIs, daemons, kernel drivers, firmware, and hardware
  • Building provisioning, test automation, and monitoring tools for lab systems.



Looking for engineers who have:

  • 8+ years of systems-software experience
  • Strong Linux systems development and low-level userspace experience
  • Strong C, plus Go, Rust, C++, or Python
  • Experience building REST APIs and CLI tools for hardware or infrastructure management
  • An understanding of device drivers, telemetry paths, PCIe devices, and BMC-managed systems
  • Experience debugging across software, firmware, and hardware boundaries
  • The autonomy to build working software against new and evolving hardware.


Nice to have:

  • Redfish, OpenBMC, gNMI, or IPMI experience
  • Cluster or fleet management for GPUs or AI accelerators
  • Hardware bring-up, lab automation, or manufacturing-test experience
  • Firmware-update, secure-boot, or attestation experience.



This is a systems role rather than a conventional web-services position. You’ll work close to the hardware while shaping how a first-generation AI platform is operated from a single node through a full cluster.

Candidates must be authorized to work in the United States and able to work from the Mountain View office Tuesday through Thursday.


Worth a confidential chat?

Anna Button Researcher

Apply for this role