layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
New York, Austin, Miami, Dallas
Source: Appnovation Technologies careers · View original posting
From Appnovation Technologies's posting. “We” and “our” refer to the employer.
Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth. Bold ambition. Practical action. Endless possibilities.
We’re looking for a Site Reliability Engineer to keep a shared observability platform for LLM-based applications running for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, with ClickHouse as the analytical store, PostgreSQL for metadata and Redis as the ingestion queue, all delivered through Argo CD.
Two things need building rather than maintaining. The platform has no monitoring, alerting or defined service levels today, and you will own putting them in place. Infrastructure is defined as code throughout, in Kubernetes manifests, Helm values and Argo CD applications.
Alongside the platform itself, you will own the runbooks that make an incident survivable by someone other than the author, and the onboarding and support path for the internal teams that depend on the platform.
ROLE RESPONSIBILITIES
Diagnose and resolve failures across ClickHouse, PostgreSQL, Redis, the ingestion workers and the Kubernetes layer beneath them, including ingest backpressure from queue depth and worker drain behaviour.
Own ClickHouse under the operator model, including Keeper quorum, replication, shard and replica topology, and S3 storage tiering.
Rehearse the restore, time it, document it and test it against its failure modes, rather than assuming a successful backup job means a recoverable system.
Plan and rehearse upgrades in a lower environment, with a rollback plan that works even if a migration is only partly complete.
Build monitoring, alerting and service levels from scratch, so problems are found here before a user reports them.
Maintain Kubernetes manifests, Helm values and Argo CD applications so every change goes through the delivery pipeline.
Write and maintain runbooks and SOPs that let a colleague resolve an incident without the author present.
Run the onboarding and support path for internal teams that depend on the platform, and triage what they bring.
Turn recurring operational work into automation.
Work with the platform engineer who owns what the platform offers. They decide what to adopt and how it is configured; you own the migration and its rollback.
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted:
True Anomaly · Denver, CO or Long Beach, CA
Supabase · Remote, Global
CWILL · CA, US; Cary, NC, US
Percepta · New York City, New York, United States
Fieldwire · United States (Remote)
Accenture Federal Services · Washington, DC