layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
Sieve
San Francisco, California, United States
Source: Sieve careers · View original posting
From Sieve's posting. “We” and “our” refer to the employer.
Sieve is a multi-modal lab curating the world's highest-quality training datasets — spanning video, audio, images, text, and 3D. We combine exabyte-scale data infrastructure and novel multimodal understanding techniques that push the frontier of foundation models. Video alone makes up 80% of internet traffic, and across modalities, data has become the enabling medium powering creativity, communication, gaming, AR/VR, and robotics.
Sieve exists to solve the biggest bottleneck in the growth of these applications: high-quality training data.
We partner with top AI labs and did $XXM last quarter alone, as a team of ~30 people. We also raised our Series A from Tier 1 firms such as Matrix Partners, Swift Ventures, Y Combinator
, and
AI Grant.
Why Now
Sieve combines access to diverse multimodal data, infrastructure to process it at scale, and close relationships with the teams building frontier models. This gives us a unique opportunity to study what makes training data effective—and turn those findings into better models and datasets.
You’ll join a small team building our research capabilities, with ownership over experiments, training systems, and the decisions those results inform.
As a Member of Technical Staff, Applied Research at
Sieve
, you’ll train and evaluate multimodal models to understand how data shapes their capabilities. Your work will span video generation and audiovisual understanding, connecting advances in data curation with measurable improvements in model performance.
You’ll own the research loop end-to-end: identify a model weakness, form a hypothesis about the data or training approach that could address it, build the experiment, and evaluate the results. This includes fine-tuning and post-training models, developing reproducible training and evaluation pipelines, and running controlled experiments on data quality, composition, and supervision.
You’re likely a good fit if you enjoy moving between research and engineering: reading a paper, implementing a method, debugging a training run, and figuring out whether an apparent improvement holds up. You care about building reliable systems and producing findings that change how we source, curate, and use training data.
Train and post-train models for video generation and multimodal understanding.
Design controlled experiments to measure how data selection, mixtures, and supervision affect model capabilities.
Build evaluations that reveal specific model weaknesses, using quantitative metrics and human judgment.
Develop reliable training infrastructure, including distributed training, efficient data loading, checkpointing, and experiment tracking.
Turn research findings into improvements in our data curation pipelines and products.
Collaborate with research and engineering teams internally and at partner labs to define meaningful problems and communicate results.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: