layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…From Upshop's posting. “We” and “our” refer to the employer.
Position Overview
We are seeking an engineering-focused
Data Scientist to build, operationalize, and maintain production-grade retail forecasting and optimization models. In this role, model development goes hand-in-hand with operational reliability: success is measured by accurate forecasts and stable, low-latency, deterministic pipelines running natively on the Databricks Lakehouse.
You will bridge the gap between applied data science and machine learning engineering. Working in close partnership with Product and
Data Engineering
, you will own the operational lifecycle of item-store level demand forecasts and downstream replenishment/optimization engines—from distributed feature pipelines to automated Databricks workflows, model tracking, and runtime monitoring.
Translate business requirements into technical specs, define operational SLAs, and provide technical feasibility assessments for new forecasting features.
Establish strict data contracts, define schema validations, and optimize data ingestion/consumption patterns from upstream Delta tables.
Treat ML pipelines as critical production software. Implement pre-inference data validation gates (e.g., schema checks, missingness thresholds, null checks) and automated alerting to prevent corrupted data from reaching scoring jobs.
Track pipeline health, monitor runtime performance, and detect feature drift, target drift, and forecast degradation across high-cardinality retail catalogs.
Write modular, maintainable, and testable code. Champion version control best practices, unit/integration testing with pytest
, and automated CI/CD checks within Git.
Embrace a Spec-Driven Development (SDD) mindset, leveraging modern agentic AI development workflows (e.g., Cursor, Claude Code) to move rapidly from research to reliable production code.
Required Technical Skills
Hands-on experience developing, tuning, and deploying gradient boosted decision trees—specifically
LightGBM and
CatBoost
—on high-cardinality, tabular, and time-series datasets.
Deep practical understanding of demand forecasting challenges: trend, seasonality, calendar events, promotional uplifts, stockouts, and intermittent/sparse demand patterns.
MLflow across the model lifecycle.
Strong proficiency in Python and PySpark for distributed data processing, feature engineering, and memory-conscious transformations across massive retail datasets.
Solid understanding of clean code principles, modular package design, virtual environments, automated testing ( pytest
), and standard Git workflows (pull requests, branching, code reviews).
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted:
Clera · San Francisco, California, United States
Clera · San Francisco, California, United States
Clera · San Francisco, California, United States
Clera · San Francisco, California, United States
Clera · San Francisco, California, United States
xai · Palo Alto, CA; Austin, TX; New York, NY; Seattle, WA