Backend Engineer - AML Engine Orchestration

3 Minutes ago • 3 Years +
Backend Development

Job Description

The Backend Engineer - AML Engine Orchestration will join ByteDance's AML team, focusing on developing next-generation machine learning algorithms and platforms for recommendation, ads, and search ranking. Key responsibilities include optimizing resource efficiency in distributed orchestration, building training system architectures for ultra-large recommendation models, and constructing online orchestration architectures for recommendation systems, driving substantial impact on core businesses.
Good To Have:
  • Familiar with large-scale distributed scheduling systems like Kubernetes, Yarn, Flink and/or Spark.
  • Familiar with opensourced orchestration frameworks like VeRL, vLLM, Ray or TFX, etc.
Must Have:
  • Develop and extend distributed orchestration frameworks within the Kubernetes/Godel ecosystem.
  • Integrate and expand AutoScaling and automatic parallelization capabilities for various models and tasks.
  • Responsible for preemption and re-scheduling mechanisms for services with different priorities.
  • Build training system architecture for next-generation ultra-large and ultra-deep recommendation models.
  • Design and optimize distributed computing APIs and runtimes geared towards future recommendation and ads model paradigms.
  • Construct robust and stable distributed model inference architecture for online learning scenarios.
  • Optimize the usability of online recommendation and ads model architectures and MLops workflows.
  • Bachelor's degree or above, majoring in Computer Science, Engineering or related fields.
  • Strong programming and coding experience with at least one modern language such as Golang, Python.
  • Experience contributing to large scale distributed systems, multi-tenant systems.
  • Strong analytical abilities and problem solving.
  • At least 3 years of relevant experience.
Perks:
  • Inspiring creativity
  • Global, diverse teams
  • Opportunity to create value for communities
  • Culture of constant iteration and "Always Day 1" mindset
  • Opportunities for growth

Add these skills to join the top 1% applicants for this job

communication
game-texts
load-balancing
spark
reinforcement-learning
yarn
kubernetes
python
algorithms
machine-learning

Responsibilities

ByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa. Team Introduction The mission of our AML team is to push next-generation machine learning algorithms and platforms for the recommendation system, ads ranking and search ranking in our company. We also drive substantial impact on core businesses of the company.

1. Resource Efficiency Optimization in Distributed Orchestration and Scheduling:

  • Develop and extend distributed orchestration frameworks within the Kubernetes/Godel ecosystem. Select appropriate frameworks based on different business scenarios, and optimize cluster utilization and load balancing strategies according to the specific characteristics of each scenario;
  • Integrate and expand AutoScaling and automatic parallelization capabilities for various models and tasks. Employ load modeling and analytic methods for different models to automatically optimize resource requests, achieving large-scale improvements in resource usage efficiency and global optimality;
  • Responsible for preemption and re-scheduling mechanisms for services with different priorities, and manage automatic resource multiplexing across different clusters and resource types; handle scheduling and load adaptation across multi-datacenter, multi-region, and multi-cloud environments.

2. Building Training System Architecture for Next-Generation Ultra-Large and Ultra-Deep Recommendation Models:

  • Develop a flexible, elastic and robust distributed training runtime focused on hyper-scaled embeddings and large-scale GPU training;
  • Design and optimize distributed computing APIs and runtimes geared towards future recommendation and ads model paradigms (e.g., reinforcement learning, fine-tuning and/or distillation);
  • Collaborate with platform teams to enhance the diagnosability and usability of distributed training systems.

3. Constructing Online Orchestration Architecture for Next-Generation Recommendation Systems:

  • Build a robust and stable distributed model inference architecture for online learning scenarios involving hyper-scaled embeddings;
  • Optimize the usability of online recommendation and ads model architectures and MLops workflows.

Qualifications

Minimum Qualifications:

  • Bachelor's degree or above, majoring in Computer Science, Engineering or related fields.
  • Strong programming and coding experience with at least one modern language such as Golang, Python.
  • Experience contributing to the large scale distributed systems, multi-tenant systems (architecture, reliability and scaling).
  • Strong analytical abilities and problem solving.
  • Good communication, self-motivation, engineering practice, documentation, etc.
  • At least 3 years of relevant experience.

Preferred Qualifications:

  • Familiar with large-scale distributed scheduling systems like Kubernetes, Yarn, Flink and/or Spark.
  • Familiar with opensourced orchestration frameworks like VeRL, vLLM, Ray or TFX, etc.

Job Information

About Us

Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Lemon8, CapCut and Pico as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.​

Why Join ByteDance

Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day.​

As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us.​

Diversity & Inclusion​

ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.​

Set alerts for more jobs like Backend Engineer - AML Engine Orchestration
Set alerts for new jobs by bytedance
Set alerts for new Backend Development jobs in Singapore
Set alerts for new jobs in Singapore
Set alerts for Backend Development (Remote) jobs

Contact Us
hello@outscal.com
Made in INDIA 💛💙