SlipstreamJobsFresh Startup & VC-Backed Jobs

Data Center Infrastructure Server Operations Expert, DCS

ByteDance - Singapore, Singapore - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ByteDance's Server Operations Team is seeking a Data Center Infrastructure Server Operations Expert to manage the full lifecycle operation of ByteDance's massive server fleet. This role combines infrastructure engineering with systems automation to deliver industry-leading server operation and maintenance services. Key responsibilities include: • Design and develop server firmware management systems and BMC (Baseboard Management Controller) management systems, building automated operation and maintenance capabilities with GUI interfaces • Architect and implement large-scale automated server operation and maintenance systems, including automated capabilities for commands, scripts, tools, and firmware upgrades • Lead the development of automated solutions for BMC management networks in collaboration with security and network engineering teams, ensuring secure and controllable operation and maintenance environments • Perform daily hardware and software operation and maintenance, online troubleshooting, fault handling, and diagnosis of complex infrastructure issues • Develop software tools and systems to ensure efficient and stable operation of automation services across the infrastructure Required qualifications include a Bachelor's degree in computer science, electronic information engineering, automation, or related field. You must be proficient in mainstream server hardware architectures (x86, ARM) and capable of independent troubleshooting and optimization. Preferred skills include strong Linux system expertise, server hardware configuration and troubleshooting, Shell scripting, and proficiency in Python or Golang. The ideal candidate demonstrates excellent problem-solving abilities, meticulous attention to detail, and the ability to handle complex system and hardware faults while developing both emergency response and long-term solutions.

Similar roles