Cloud Native Engineer, ARK Large Model Platform (Singapore)

3 Months ago • 3 Years + • Research & Development • Artificial Intelligence

Job Summary

Job Description

ByteDance is seeking a Cloud Native Engineer to join their Applied Machine Learning (AML) - Enterprise team in Singapore. The role will focus on developing and maintaining the ARK Large Model Platform on VolcanoEngine, a cloud native resource scheduling system. The ideal candidate will have strong experience in cloud computing, large-scale model systems, and Golang/C++/Cuda development. The position involves tackling challenging tasks related to large language models, distributed training, and cluster management. Responsibilities include building efficient training and inference systems, managing distributed training jobs, constructing scalable ML systems, and investigating cutting-edge technologies in AI and machine learning.
Must have:
  • B. Sc or higher degree in Computer Science or related fields
  • 3+ years of R&D experience in cloud computing or large-scale model systems
  • Experience in Golang/C++/Cuda development
  • Understanding of Linux systems and cloud platforms
  • Knowledge of cloud-native orchestration technologies like Kubernetes
  • Experience in large-scale cluster maintenance and optimization
  • Grasp of computer networking, Linux file system, object storage services, SQL and NoSQL databases
  • Self-motivated, innovative, collaborative, and uphold high coding and documentation standards
Good to have:
  • Experience in developing ML platforms or MLOps platforms
  • Experience in distributed machine learning model training, ML model fine-tuning, and deployment

Job Details

Responsibilities
ByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa. Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content. Why Join Us Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. To us, every challenge, no matter how ambiguous, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. Join us. About the Team The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services. In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities. Responsibilities Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world. - Maintain a large-scale AI cluster and develop state-of-the-art machine learning platforms to support a diverse group of stakeholders. - Tackle extremely challenging tasks which include, but are not limited to, delivering highly efficient training and inference for large language models, managing extremely effective distributed training jobs across clusters with over 10,000 nodes and GPU chips, and constructing highly reliable ML systems with unparalleled scalability. - The work encompasses various aspects of LLMOps (Large Language Model Operations), such as resource scheduling, task orchestration, model training, model inference, model management, dataset management, and workflow orchestration. - Investigate cutting-edge technologies related to large language models, AI, and machine learning at large, such as state-of-the-art distributed training systems with heterogeneous hardware, GPU utilization optimization, and the latest in hardware architecture. - Employ a variety of technological and mathematical analyses to enhance cluster efficiency and performance.
Qualifications
Minimum Qualifications - B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions with at least 3 years of R&D experience in the fields of cloud computing or large-scale model systems. - Experience in Golang/C++/Cuda development with a solid understanding of Linux systems and popular cloud platforms such as Volcano Engine Cloud, AWS, and Azure Cloud. - Profound knowledge of cloud-native orchestration technologies like Kubernetes, coupled with experience in large-scale cluster maintenance, job scheduling optimization, and cluster efficiency enhancement. - A strong grasp on various foundational areas of computer science, including computer networking, the Linux file system, object storage services, and SQL as well as NoSQL databases. - Self-motivated, thirst for innovation, collaborative working aptitude, and consistently uphold high standards in coding and documentation quality. Preferred Qualifications: - Experience in developing ML platforms or MLOps platforms. Experience in distributed machine learning model training, ML model fine-tuning, and deployment. ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.

Similar Jobs

31st Union - Expert Core Engineer

31st Union

San Mateo, California, United States (On-Site)
• 1 Month ago
Microsoft - Software Engineer 2 (Core Search platform)

Microsoft

Beijing, Beijing, China (On-Site)
• 1 Month ago
NXP - Software Engineering Intern, Linux Kernel/BSP

NXP

Bucharest, Bucharest, Romania (On-Site)
• 5 Months ago
ByteDance - Site Reliability Engineering, Edge Services - Traffic Infrastructure

ByteDance

Singapore (On-Site)
• 3 Months ago
Tama Systems India   - Embedded Engineer

Tama Systems India

Bengaluru, Karnataka, India (On-Site)
• 4 Months ago
Krafton  - PUBG IP Franchise Project ARC Community Manager

Krafton

Seoul, South Korea (On-Site)
• 1 Month ago
Cirrus Logic - Manager, Design Engineering (MMS-64000105)

Cirrus Logic

Edinburgh, Scotland, United Kingdom (Hybrid)
• 4 Months ago
NVIDIA - Senior CPU and SOC Verification Engineer

NVIDIA

Bengaluru, Karnataka, India (On-Site)
• 1 Month ago
Patterned Learning Career - Vice President, Software Development, Automation, Material Handling

Patterned Learning Career

(Remote)
• 1 Week ago
NVIDIA - Physical Design Backend Engineer

NVIDIA

Yokne'am Illit, North District, Israel (Hybrid)
• 1 Week ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Google - Software Engineer II, Google Cloud

Google

(On-Site)
• 2 Months ago
Riot Games - Principal Software Engineer, Product Tech-Lead - Unpublished R&D Product

Riot Games

Dublin, County Dublin, Ireland (On-Site)
• 3 Months ago
Mojang Studios - Senior C++ Gameplay Engineer (Bedrock)

Mojang Studios

Stockholm, Stockholm County, Sweden (On-Site)
• 5 Months ago
Riot Games - Staff Software Engineer, Services - MMO

Riot Games

Los Angeles, California, United States (On-Site)
• 2 Weeks ago
Sony Interactive Entertainment - Application Security Engineer

Sony Interactive Entertainment

Tokyo, Japan (On-Site)
• 1 Month ago
Microsoft - Software Engineer

Microsoft

Noida, Uttar Pradesh, India (On-Site)
• 4 Weeks ago
ByteDance - Senior Software Engineer, Cross Platform

ByteDance

San Jose, California, United States (On-Site)
• 3 Months ago
Guerrilla - Lead UI Programmer

Guerrilla

Amsterdam, North Holland, Netherlands (On-Site)
• 1 Month ago
Build A Rocket Boy - Technical Artist

Build A Rocket Boy

Edinburgh, Scotland, United Kingdom (On-Site)
• 1 Month ago
PhonePe - PSE - Data Engineering

PhonePe

Bengaluru, Karnataka, India (On-Site)
• 3 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Singapore

ByteDance - Principal Site Reliability Engineer, CDN

ByteDance

Singapore (On-Site)
• 3 Months ago
ByteDance - Technical Support Engineer, Video Cloud

ByteDance

Singapore (On-Site)
• 3 Months ago
ByteDance - LLM Performance Operation, Analysts (Safety)

ByteDance

Singapore (On-Site)
• 2 Months ago
ByteDance - Backend Engineer - Applied Machine Learning Platform

ByteDance

Singapore (On-Site)
• 3 Months ago
Bushiroad - Sales Executive

Bushiroad

Singapore, Singapore (On-Site)
• 2 Weeks ago
NAH.io - Vice President of Human Resources (Hedge Fund)

NAH.io

Singapore (Hybrid)
• 3 Months ago
The Walt Disney Company - MarkOps Consultant - Contract

The Walt Disney Company

Singapore, Singapore (On-Site)
• 3 Months ago
Netflix - Associate, Revenue Analytics

Netflix

Singapore, Singapore (On-Site)
• 1 Month ago
Interactive Brokers - Associate - Client Services

Interactive Brokers

Singapore (Hybrid)
• 4 Months ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

Netflix - Manager, Security Protocols Engineering

Netflix

United States (Remote)
• 3 Months ago
NVIDIA - Senior Chip Design Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (Hybrid)
• 1 Week ago
NVIDIA - Senior ASIC Floorplan Design Engineer

NVIDIA

Santa Clara, California, United States (Remote)
• 1 Month ago
NVIDIA - Senior AI and HPC Modeling Architect

NVIDIA

Yokne'am Illit, North District, Israel (On-Site)
• 1 Month ago
Cirrus Logic - Summer Intern, Analog Design Engineer

Cirrus Logic

Austin, Texas, United States (On-Site)
• 4 Months ago
NVIDIA - Senior Software Engineer - System Customization Team

NVIDIA

Yokne'am Illit, North District, Israel (On-Site)
• 1 Month ago
Microsoft - Research Intern - Machine Learning for Biology and Healthcare

Microsoft

Cambridge, Massachusetts, United States (On-Site)
• 1 Month ago
Advanced Sterilization Products - Senior Software Engineer - Java Fullstack

Advanced Sterilization Products

Bengaluru, Karnataka, India (Hybrid)
• 4 Months ago
Cadence - Lead Design Engineer ( Layout Design )

Cadence

Bengaluru, Karnataka, India (On-Site)
• 4 Months ago
NVIDIA - Research Scientist, Network - New College Grad 2025

NVIDIA

Santa Clara, California, United States (On-Site)
• 1 Month ago

Get notifed when new similar jobs are uploaded

About The Company

Where imagination meets innovation, delivering limitless gaming experiences.

Taguig, Metro Manila, Philippines (On-Site)

Singapore (On-Site)

Dubai, Dubai, United Arab Emirates (On-Site)

State Of São Paulo, Brazil (On-Site)

Seattle, Washington, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by ByteDance

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug