Opplane is applying AI across the software delivery lifecycle - not only to writing code, but to testing, review, documentation, migration, incident response, and the validation stages where delivery is usually constrained. Measuring what that produces is the starting project, not the whole of it.
You will own the analytical and applied half of that work: designing the experiments that establish what actually helps, building the classification and evaluation pipelines the measurement program depends on, and then taking the findings back into the delivery lifecycle as capabilities teams can use.
This role pairs with a data engineer who owns extraction, identity, and the metric pipeline. You own what the numbers mean and what to do about them.
Design and run experiments. Model tier routing, MCP coverage, permission configuration, repository context quality, budget headroom — randomized across teams and reported with their limits stated. These are the cleanly causal questions available once a tool is deployed, and where the returns are.
Own the analytical layer of the measurement program: work classification over model traffic, evaluation design, longitudinal within-unit analysis, and the staggered-adoption estimates that connect delivery outcomes to adoption timing.
Build and validate LLM-as-judge and classification pipelines — sampling strategy, hand-labeled ground truth, precision and recall measured and published, and revalidation whenever the taxonomy or the model changes.
Extend AI beyond code authoring into the stages that constrain delivery: test authoring and maintenance, environment and data setup, migration and modernization, code review assistance, security remediation, and evidence assembly for certification.
Work directly with the constrained teams. Where validation or certification is the bottleneck, coding assistance produces little regardless of how well it works — find where the constraint actually sits and aim the capability at it.
Build evaluations for internal AI capabilities: golden sets, regression suites, groundedness and answer-quality scoring, and the cost and latency telemetry beside them.
Turn findings into practice. Identify what the most effective practitioners do differently, document it, and teach it — publishing practices rather than rankings.
Partner with the platform team on Claude Code configuration, MCP servers, gateway telemetry, and the model registry, so what you learn becomes the default rather than folklore.
Report to engineering leadership and finance: what is working, what is not, where delivery is actually constrained, and what each claim does and does not establish.
7+ years spanning software engineering and quantitative analysis. This role needs both; a strong background in one and a passing acquaintance with the other will not carry it.
Production experience with LLM applications: prompting, tool and function calling, context management, evaluation, and knowing where models fail in practice.
Experimental design and causal inference — randomized and quasi-experimental designs, difference-in-differences, instrumental variables, hierarchical models — and the judgment to say when a design does not support the claim being asked of it.
Strong Python and SQL, with a statistical stack (pandas, statsmodels, scikit-learn, or R). Data collection: instrumenting and extracting from operational systems and APIs, designing sampling that survives scrutiny, and knowing when a source cannot answer the question being asked of it.
Aggregation: resolving identity across systems, joining sources never designed to be joined, and modeling the summary tables reporting reads from. You need not own the pipeline, but you must be able to build one when the answer depends on it.
Analytics: exploratory analysis, distributions rather than averages, cohort and time-series work, and reports that state their own coverage and limits.
Real familiarity with the software delivery lifecycle — code review, CI/CD, test strategy, release and change management — sufficient to hold a credible conversation with the teams you are measuring.
Git and GitLab at instrumentation depth: merge request and pipeline data models, diffs and SHAs, what merge, squash, rebase, and cherry-pick do to line-level analysis, and the API and hook surfaces available for capturing it.
Jira and Confluence integration experience — REST APIs, changelog and page version history, the GitLab–Jira development panel, and the field and label conventions that determine whether the resulting data means anything.
Care with personnel-adjacent data: aggregate reporting by default, and a clear sense of what should not be built even when it is technically easy.
Communication that works in both directions — an executive audience that wants a number, and engineers who will dispute it.
MCP servers and clients, or comparable connector frameworks.
Agent frameworks — LangGraph, LangChain, Bedrock Agents, Strands, or equivalent.
Enterprise deployment of coding assistants, and their telemetry.
Server-side Git hooks, GitLab CI, and system or webhook-driven capture on a self-managed instance.
Confluence and Jira as MCP-connected systems — permission propagation, scoped credentials, and audit logging. Evaluation tooling — Ragas, DeepEval, Bedrock model evaluation — and LLM observability such as LangFuse, Arize, or OpenTelemetry-based tracing.
Amazon Bedrock, and AWS cost and usage data.
Engineering productivity frameworks — DORA, DX Core 4, SPACE — and a working view of their limits.
Program analysis, test generation, or developer tooling research.
dbt, Airflow, Dagster, or equivalent transformation and orchestration; warehouse or lakehouse modeling.
BI and visualization tooling, and the discipline of building on summary tables rather than raw events.
Queueing and flow analysis: utilization, batch economics, and constraint identification.
Opplane specializes in providing advanced data-focused solutions for financial services, telecommunication, and reg-tech to accelerate their digital transformation journey. Opplane leadership team is comprised of Silicon Valley serial entrepreneurs and experienced executives. Its expertise comes from years of specific industry experience at some of the world’s top companies, such as PayPal, Xerox Parc, Amazon, Wells Fargo, SoFi in the areas of product management, data technology, data governance, data privacy, security, machine learning, and risk management.
🌍 Global & Multicultural – Diverse perspectives, global collaboration (US, Portugal, India and Singapore offices)
⚡ Startup Energy – Fast-moving, impact-driven environment
💪 Ownership Mindset – Engineers own what they build
🤝 Collaborative & Friendly – Open, curious, and supportive culture
Most organizations deploying AI to engineering cannot say what they got for it and respond by buying more of it or by arguing. You will build the evidence instead and then use it to decide where the next capability goes.
The measurement program is the first project. The scope is applied AI across the delivery lifecycle, with real internal users, a platform team to build on, and leadership that will act on what you find.