SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cloudflare's Infrastructure group is responsible for building and maintaining the global network that powers 20% of the world's internet traffic across 330 cities in 120+ countries. As a Hardware Systems Engineer, you will be part of the Hardware Engineering team working to research, develop, test, and deploy equipment that improves the security, reliability, and performance of the Internet.
You will work collaboratively with Hardware Engineering, Product, and Hardware Sourcing teams to troubleshoot and maintain Cloudflare's worldwide fleet of storage and compute servers. Your responsibilities will include validating bug fixes and assessing performance of new firmware revisions, deploying firmware updates to the fleet while monitoring rollout compliance and reliability, and working with server and component vendors to obtain, debug, and maintain the latest updates.
You will support Site Reliability Engineering teams in triaging hardware problem reports and collaborate with Data Centre Engineering teams to resolve hardware issues. A key part of the role involves developing and maintaining automation tools to update firmware on servers and components across Cloudflare's fleet. You will also communicate your results and updates through blog posts, internal talks, and tickets.
The hardware you work with includes servers and components, PDUs, and network hardware. This role requires understanding how a diverse server fleet is managed at scale and the tools used to maintain and monitor hardware in a production mission-critical environment.
**Requirements:**
- 3-5 years of experience
- Bachelor's degree in Computer Engineering, Electrical Engineering, or Computer Science
- Knowledge of bash and python and basic Linux task automation
- Knowledge of x86 server hardware including motherboards, CPUs, memory, storage, and firmware updates
- Knowledge of configuration management principles (experience with salt preferred)
- Knowledge of Redfish, IPMI, and server remote management protocols
- Experience running production mission-critical systems
**Desirable Skills:**
- Knowledge of other platforms such as ARM
- Familiarity with server hardware architecture
- Knowledge of debugging server hardware faults and vendor engagement
- Experience managing large fleets of thousands of servers
- Experience with observability and monitoring tools such as Prometheus and Grafana
- Experience with software development tools and processes (git, Bitbucket, TeamCity, Jira)