layiq
worthy; deserving; fitting; suitable.
A role, opportunity, or path that merits attention, time, and pursuit.
Loading LAYIQ…Job opportunity
Microsoft
Redmond, WA
Source: Microsoft careers · View original posting
From Microsoft's posting. “We” and “our” refer to the employer.
Overview
We are building the evaluation backbone for safe, reliable, and efficient agentic engineering in Microsoft Security. This team will create systems that determine when an AI agent, model, prompt, tool, memory strategy, or orchestration pattern is ready to be used in production security and engineering workflows.
The role is ideal for engineers who can operate end to end: understand the workflow, design the benchmark, build the harness, implement validators and graders, run experiments, analyze quality and cost tradeoffs, connect results to production feedback, and help teams make evidence-based release decisions.
Why this role matters
: Microsoft Security is moving toward agentic engineering systems for security triage, remediation, repo readiness, and scan-to-verified-closure workflows. Evals are the trust system for that shift. They help decide whether autonomy can safely expand, whether a release should stop, and which configuration achieves the required quality, safety, reliability, latency, and cost bar with the lowest practical human-review burden.
Role mission
As a Principal Security Research Manager on the AI Evaluation Systems team, you will build the common evaluation platform and methodology used by MSec agent programs. You will work across evaluation design, platform implementation, test infrastructure, telemetry, measurement, security workflow understanding, and production learning. Your work will make agentic systems measurable, reproducible, governable, and continuously improving.
Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Versioned benchmark suites for security triage, remediation, repo readiness, tool use, escalation, and end-to-end agentic workflows.
Evaluation runners and harnesses that can replay tasks, capture traces, evaluate outputs, and compare multiple agent configurations.
Deterministic validation checks for code, policy, security, provenance, ownership, deployment constraints, and workflow-specific correctness.
Automated and human-in-the-loop grading pipelines with calibration, sampling, and confidence thresholds.
Pareto-style scorecards that show tradeoffs across quality, risk, latency, tokens, cost, and human-review burden.
Telemetry and production feedback loops that continuously expand benchmark coverage and keep offline evaluation anchored to real-world outcomes.
Release-readiness gates and evidence packages that help leaders and product teams decide whether to scale, stop, or redesign an agent pattern.
Security Research M5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here: https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable
LAYIQ is an independent job-discovery service. This listing does not imply a partnership with or endorsement by the employer. Review the original posting for current details and availability.
Employer posted: