Big Data Engineer, Data Lake / Feature Store

6 Months ago • 2 Years + • Monetization

Job Summary

Job Description

The batch processing team at ByteDance is responsible for offline data processing and distributed training. You will be developing and optimizing the in-house Feature Store functionality based on Iceberg, participating in optimizing the integration of Iceberg with various upper-level computing engines, and being involved in platform-related infrastructure development.
Must have:
  • Bachelor's Degree or above in Computer Science or related fields
  • 2+ years of relevant development experience
  • Strong programming ability in Java, Python, C++
  • Experience with large-scale distributed systems
  • In-depth knowledge of data lake formats like Delta, Hudi, or Iceberg
Good to have:
  • In-depth research or practical experience in Hadoop, Spark, Flink, Presto
  • Experience with open-source big data computing frameworks

Job Details

Responsibilities
About ByteDance Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content. Why Join Us Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible. Together, we inspire creativity and enrich life - a mission we aim towards achieving every day. To us, every challenge, no matter how ambiguous, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always. At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve. Join us. About The Team The batch processing team is responsible for the company's offline data processing and distributed training, supporting various business scenarios such as offline ETL and machine learning within the company. The components involved include the offline computing engine Spark, the in-house distributed training framework Primus, feature storage solutions like Iceberg and Hudi, as well as Ray, a next-generation distributed application framework. Faced with massive-scale scenarios, extensive functional and performance optimizations have been carried out in Spark, Primus, Feature Store, and support for the adoption of the new-generation distributed application framework Ray in relevant company scenarios. What you will be doing: - Responsible for the development and performance optimisation of the in-house Feature Store functionality based on Iceberg; - Participant in optimisation of the integration of Iceberg with various upper-level computing engines; - Involve in platform-related infrastructure development.
Qualifications
Minimum Qualifications - Bachelor's Degree or above, majoring in Computer Science, or related fields, with 2+ years of relevant development experience in the field with a strong programming ability, and proficiency in Java, Python, C++, with the ability to develop and optimize large-scale distributed systems. - In-depth research and relevant experience in one or more data lake formats such as Delta, Hudi, or Iceberg. Preferred Qualifications - In-depth research or practical experience in open-source big data computing frameworks and scenarios like Hadoop, Spark, Flink, Presto, and more. ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too. #LI-CT

Similar Jobs

Glean - Solutions Architect ( EMEA/US East Customer hours )

Glean

Bengaluru, Karnataka, India (On-Site)
5 Months ago
STAGE - Analytics Engineer

STAGE

Noida, Uttar Pradesh, India (On-Site)
8 Months ago
Google - Software Engineering Manager, Cloud AI

Google

Warsaw, Masovian Voivodeship, Poland (On-Site)
2 Weeks ago
Postman - Senior Backend Engineer, Cloud Platform

Postman

San Francisco, California, United States (On-Site)
6 Months ago
Google - Staff Software Engineer, Google Cloud Networking

Google

Warsaw, Masovian Voivodeship, Poland (On-Site)
2 Weeks ago
ByteDance - SoC System Software Architect

ByteDance

San Jose, California, United States (On-Site)
6 Months ago
ByteDance - Machine Learning Scientist Graduate, Scaling AI for Biology (AML - AI-for-Science) - 2025 Start (PhD)

ByteDance

Seattle, Washington, United States (On-Site)
6 Months ago
Playtika - Technical Operation Specialist

Playtika

Israel (On-Site)
6 Months ago
Voodoo - Monetization Associate - Blitz

Voodoo

Paris, Île-de-France, France (On-Site)
3 Months ago
ByteDance - Transfer Pricing Manager - US

ByteDance

San Jose, California, United States (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Next Level Business Services - Java/J2EE Developer

Next Level Business Services

Tampa, Florida, United States (On-Site)
6 Months ago
Anavation - Senior Software Engineer

Anavation

Huntsville, Alabama, United States (Remote)
1 Week ago
Google - Software Engineer, Early Career

Google

Bucharest, Bucharest, Romania (On-Site)
2 Days ago
Google - Software Engineering Manager II, Infrastructure, Google Cloud Compute

Google

Seattle, Washington, United States (On-Site)
2 Weeks ago
Velotio Technologies - Lead Engineer (Java)

Velotio Technologies

Pune, Maharashtra, India (Remote)
1 Month ago
NVIDIA - Staff Integration Engineer

NVIDIA

Santa Clara, California, United States (Hybrid)
6 Days ago
Google - Software Engineer II, Cloud Networking

Google

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Weeks ago
ByteDance - Software Engineer - Serverless Compute Infrastructure

ByteDance

San Jose, California, United States (On-Site)
2 Months ago
Hashlist - Senior Data Engineer

Hashlist

Pune, Maharashtra, India (Hybrid)
5 Months ago
Canva - Software Engineer Internship (Infrastructure)

Canva

Auckland, Auckland, New Zealand (Remote)
3 Weeks ago

Get notifed when new similar jobs are uploaded

Jobs in Singapore

ByteDance - Security Engineer, Security Assurance

ByteDance

Singapore (On-Site)
1 Month ago
ByteDance - Product Design Intern - Global Payment

ByteDance

Singapore (On-Site)
1 Month ago
Google - Strategic Partner Manager, Tier 1 Apps

Google

Singapore, Singapore (On-Site)
2 Days ago
The Walt Disney Company - Creative Manager, IM SEA

The Walt Disney Company

Singapore, Singapore (On-Site)
5 Months ago
Google - Social Insight Strategist

Google

Singapore (On-Site)
2 Days ago
ByteDance - Security Software Engineer

ByteDance

Singapore (On-Site)
2 Weeks ago
ByteDance - PMO Team Lead

ByteDance

Singapore (On-Site)
1 Month ago
HoYoverse - CRM Lifecycle Manager

HoYoverse

Singapore, Singapore (On-Site)
19 Hours ago
ByteDance - Software Engineer Intern, Security Engineering

ByteDance

Singapore (On-Site)
1 Month ago
ByteDance - Software Engineer - Low-code Platform

ByteDance

Singapore (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Monetization Jobs

Sunday - Growth Manager

Sunday

Hamburg, Hamburg, Germany (Hybrid)
1 Week ago
InMobiInMobi - Director, Publisher Development

InMobiInMobi

London, England, United Kingdom (On-Site)
1 Month ago
ByteDance - Site Reliability Engineer - AML

ByteDance

San Jose, California, United States (On-Site)
6 Months ago
Playtika - Monetization Manager

Playtika

Israel (On-Site)
6 Months ago
ByteDance - Brand Partnership Manager

ByteDance

Austin, Texas, United States (On-Site)
2 Weeks ago
Joyteractive - Segmentation Producer

Joyteractive

Georgia (Remote)
1 Month ago
Anzuio - Senior Account Executive

Anzuio

Massachusetts, United States (Hybrid)
1 Month ago
Netflix - Manager, FS&A Ads (Platform)

Netflix

Los Angeles, California, United States (On-Site)
6 Months ago
Inwave - Monetization Specialist

Inwave

(On-Site)
2 Months ago
ByteDance - Lead Research Scientist, Foundation Model, Music Intelligence

ByteDance

San Jose, California, United States (On-Site)
6 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Where imagination meets innovation, delivering limitless gaming experiences.

San Diego, California, United States (On-Site)

San Jose, California, United States (On-Site)

Dubai, Dubai, United Arab Emirates (On-Site)

New York, New York, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

Seattle, Washington, United States (On-Site)

View All Jobs

Get notified when new jobs are added by ByteDance

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug