NEXUSTalent Partners Find a Job
Home › Jobs › I.T. & Communications › Platform Engineer

Platform Engineer

GCB Services LLC
Sunnyvale, California I.T. & Communications Today
Apply Now

Job description

Role 1 - Core Platform Engineer (L1, breadth-first) The engineering first line of defense: incident response, triage, reliability, and automation across the full infrastructure stack. Day-to-day: Runs incident response drills, post-mortems, and root cause analysis; learns from past incidents to prevent recurrence. Starts the day reviewing overnight alerts and system performance metrics, triaging anomalies. Participates in team stand-ups on projects, incidents, and daily priorities. Automates routine processes, analyzes system logs, and builds tools to strengthen monitoring. Works alongside software engineers advising on resilient-code best practices and reviewing changes pre-deployment. Maintains high SLIs/SLOs; documents work and shares insights with a customer-centric mindset. Must have: Architecture, design patterns, reliability, and scaling of new and existing systems. Incident command experience - driving RCA, coordinating cross-functional teams, ensuring corrective-action follow-through. Observability built from the ground up - defining SLOs/SLIs, closing monitoring gaps, alerting strategies that catch failures before customers do. Linux kernel internals - scheduler, memory allocation, driver subsystems. High-quality code in at least one language (Python, Go, or similar). System-level debugging - kdump, kernel panic analysis. IaC (Ansible, Terraform, Kubernetes) and CI/CD (GitLab CI, AWX, etc.) for bare-metal or cloud infrastructure. TCP/IP and network programming. Distributed storage systems - object, block, and/or file storage paradigms. Strong communication skills. Nice to have: Hardware and GPU troubleshooting. OVN/OVS-based networking stack exposure. Sourcing note: This is a deep SRE profile, not a pure generalist. The kernel-internals and system-level debugging bar is real and higher than a typical "L1" label implies - screen for genuine engineering depth, not helpdesk/NOC-tier breadth. Role 2 - Platform Engineer (L2, depth-first) Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work. Must have: Deep expertise in one domain (Storage / Compute / SDN). SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain. Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs. Domain specifics: o Storage: block, blob, and file storage; distributed storage; performance diagnostics and data-path optimization. Critical screening filter - operator vs. SRE: Crusoe explicitly does not want another "operator" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure. Role 3 - Platform Engineer (L2, depth-first) Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work. Must have: Deep expertise in one domain (Storage / Compute / SDN). SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain. Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs. Domain specifics: o Compute: Linux systems, KVM/QEMU, Cloud Hypervisor, kernel tuning, CPU/memory/VM optimization. Critical screening filter - operator vs. SRE: Crusoe explicitly does not want another "operator" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure. Role 4 - Platform Engineer (L2, depth-first) Specialized domain expert embedded in a single foundation team: Storage, Compute, or SDN. Owns reliability, performance, scalability, and operational excellence for that one pillar. No cross-vertical work. Must have: Deep expertise in one domain (Storage / Compute / SDN). SRE as a fundamental minimum - has set up observability, done alert management, and worked with SLIs/SLOs within that domain. Ability to code the infrastructure: extend out-of-the-box product parameters to build observability layers, and tie SLAs/SLOs to business KPIs. Domain specifics: SDN: OVS/OVN, network virtualization, high-performance networking, NIC tuning. Critical screening filter - operator vs. SRE: Crusoe explicitly does not want another "operator" (storage admin doing patching, installs, upgrades). Screen hard for SRE substance (observability built, alerting, SLI/SLO ownership), not just domain tenure.
Apply for this job

You'll be taken to the employer's application page.

Similar jobs

Software Engineer III
Pinnacle Technical Resources
Sunnyvale
Today
Data Engineer II
Pinnacle Technical Resources
Sunnyvale
Today
Test software engineer(LabVIEW, Matlab, Python, and C#)_Sunnyvale, CA
Trispark Inc.
Sunnyvale
Today
Software Engineer
Aditi Consulting
Sunnyvale
Today
Product Operations Manager
CYNET SYSTEMS
Sunnyvale
Today
Test Software Engineer (LabVIEW, Matlab, Python)
Tanisha Systems
Sunnyvale
Today