Building agent systems · Open to Staff & Principal, or SDM & EM

Dhiraj Sharma

Senior Software Engineer, Microsoft AI — Bing Ads, Data Platform as a Service

I build systems that coordinate work at scale — fleets of AI agents that decompose, fan out, and verify before they finish.

  • 6M/minevents at 99.99%+
  • $14Minfra removed
  • 38K+/dayHadoop job runs
  • 3,400Hadoop cores
  • 48%faster batch
  • 10+years shipping

Armada

Multi-agent orchestration · Microsoft AI, productized independently

One agent is an assistant. A fleet is a force multiplier.

Ran a real Azure migration with up to eight agents in parallel — minutes, not a night — while the human only reviewed.

Watch the workflow — 1:48one sequential agent versus a fleet, on a real Azure migration: dependency classification, eager fan-out, warm reuse, and up to eight agents in parallel while the human only reviews.
A task decomposes, fans out to agents, and those agents fan out again — nesting as deep as the work needs. Reconciliation is dynamic and runs across fans, not just up the tree: sibling branches settle with each other as results land, then fold level by level, one branch owning any shared resource, until a single result passes the verification gate. Shipped on GitHub Copilot, Copilot CLI, and VS Code from one canonical core.
Inside the runtime
Inside: the scheduler decomposes and right-sizes each task, a lock table guarantees only one agent owns a shared resource at a time, and agents are claimed from a warm pool under a lease — reused rather than cold-started, capped so idle capacity cannot pile up, and reaped on TTL when a lease goes stale. The reconciler merges what returns, and nothing reaches "complete" without passing the gate. Work is modelled explicitly — tasks, dependencies, owners, resource claims, status and verification evidence — so a run in flight stays inspectable and recoverable rather than opaque, and partial sub-agent output can never be reported as finished work.

AI systems I've shipped

Platform toolkit 28 Terraform / CDKTF skills, each turning a repeatable platform chore into a governed self-service workflow — built on Terraform, CDKTF, Bicep, Azure DevOps and the Azure CLI, against AKS, ACR, Key Vault, Workload Identity and the Secrets Store CSI Driver, in TypeScript, Node.js and PowerShell.

  • Armada Fan-Out
  • Terraform Security Baseline Applier
  • State Drift Resolver
  • Plan Reviewer
  • Terraform PR Quality Evaluator
  • CDKTF PR Quality Evaluator
  • State Migrator
  • State Backend Configurator
  • Terraform State Storage NSP Applier
  • Terraform Provider Version Policy
  • Terraform Provider Upgrade
  • Terraform-to-CDKTF Converter
  • Terraform-to-Modules Migrator
  • CDKTF Module Binding Generator
  • CDKTF Project Bootstrapper
  • New Module Scaffolder
  • Module Usage Generator
  • Module Version Pinning
  • Module Catalog Lookup
  • Cost Impact Estimator
  • Environment Promotion Scaffolder
  • Naming Convention Applier
  • Tagging Standard Applier
  • Azure VMSS Org Standards
  • ADO Service Connection Creator
  • Providers JSON Generator
  • Terraform Error Troubleshooter
  • Skill Test Harness Scaffolder

Experience

  1. MicrosoftJun 2026 — now

    Senior Software Engineer, Microsoft AI. Joined an in-flight on-prem → Azure program; what I have owned since: Armada, Fleet Hub, the Azure networking layer and Private Link, Vault → Key Vault with drift reconciliation, and AKS job execution on ephemeral per-job VMs.

  2. Amazon2019 — 2026

    Software Development Engineer. Tier-1 push notifications at ~6M events/min, 99.99%+ availability. Alexa in-app purchase 0 → ~$4M/yr. Hey Disney integration. Add to Delivery notifications for Prime members. ~$14M/yr cut from the Rekognition fleet.

    Built the team's client-side metrics and alarming from zero — replacing hand-maintained monitors with generated, version-controlled ones — cut cross-service call volume by half, contributed components to a shared library adopted by 10+ teams, and led Prime Day and Q4 peak readiness. Became the only certified security reviewer in the org while still at SDE1.

  3. OOracle2017 — 2019

    Senior Software Engineer. Java/Spring Boot microservices on the Advanced Support Platform; production errors down ~45% via JVM profiling and query tuning. Technical lead for 10+ engineers. Product docs

  4. Care.com2016 — 2017

    Senior Software Engineer. Acquisition funnels and instrumentation — click-through +25%, conversion +21%.

  5. Samsung2016

    Senior Software Engineer. Artik marketplace features and partner portals — ~27% revenue growth.

  6. A&MTexas A&M–Kingsville2014 — 2015

    Research Assistant on the Software Visualization project — a web IDE rendering static structure as UML and dynamic behaviour from runtime stack traces in D3 — and Graduate Teaching Assistant for Operating Systems and Mobile Application Development, where I built the automated attendance and grading system the courses ran on.

  7. LLocuz Enterprise Solutions2013 — 2014

    Software/Cloud Engineer, Innovative Development Team. A cross-vendor cloud management API — configuration, installation, reporting and analysis — validated in EMC² VSPEX and NetApp FlexPod labs; a cluster manager with job submission and tracking that worked independently of the underlying cluster suite; and aCube, centralised licence management over a FlexLM server.

  8. IIT Bombay2011 — 2013

    Virtual Labs, Chemical Engineering. Built the Healthcare Research Consortium platform and the pre/post-viva modules, ran the lab's NAS and Linux compute server, and taught FOSS workshops for the Spoken Tutorial project under MHRD's National Mission on Education through ICT.

Beyond the codebase

Where the work reaches people who do not report to me.

Leading

Authority, standards and the times somebody had to pick it up.

  1. Authority to stop another team's launch

    Amazon Security Certifier from 2020 — and the only certified security reviewer in the org while still at SDE1. The role carries design-review authority over other teams' production readiness, which means the power to say a launch is not going out. Exercised across multiple services, more than the bar required.

  2. Picked it up when the team halved

    After roughly 60% of the team left in one quarter, I took Prime Day and Q4 peak readiness onto myself rather than waiting for it to be assigned, and carried the component work the wider org was depending on. Nobody appointed me to either.

  3. Changed the plan by writing the argument down

    Proposed splitting a service the team had not planned to split, wrote the design documents for both halves in parallel, and took them through review with the data behind the trade-offs. My manager was not initially convinced. The design went in as I argued it, and the second service exists because of that document.

  4. Standards other teams pick up

    Reusable service and infrastructure templates adopted across teams; components contributed to a shared library used by 10+ teams; a 28-skill platform toolkit that turns repeatable chores into governed workflows; and a persistent engineering-standards system so a team's operating rules survive between sessions rather than living in someone's head.

  5. Engineers, directly

    Technical lead for a 10+ engineer team at Oracle — design reviews, code reviews, performance debugging. Mentoring and cross-team design reviews at Amazon. At Microsoft I own each team's move onto the platform end to end: discovery, scoping, design review, hands-on migration of their workloads, and cutover support.

  6. Teaching, before any of it

    40+ open-source workshops taught across colleges in India for IIT Bombay's Spoken Tutorial project under a Government of India programme, with standing authorization to run them anywhere in the country — and a Graduate Teaching Assistant post at Texas A&M–Kingsville for Operating Systems and Mobile Application Development.

Also built

Architecture I designed

Systems I owned end to end — drawn as they actually behave, including the failure paths.

Tier-1 push notification platform

The platform behind notifications like Add to Delivery, which tells a Prime member their order can still be changed — so it is only useful if it arrives before the van does. Producers fan into SNS, SQS absorbs the spike as back-pressure, Fargate workers drain it, and every write is idempotent so a replay cannot double-send. What fails retries; what keeps failing lands in the DLQ instead of blocking the queue.

Add to Delivery — About Amazon

Data Platform as a Service on Azure — Bing Ads

The multi-region Hadoop and Spark platform where all bidding data lands, reached over Private Link — sitting on a dependency map of batch compute, an LDAP auth bridge, a MySQL metadata store, object storage, message brokers, Redis, ZooKeeper and the operational log pipelines — where single tenants run to petabytes, 1.05 PB live against a 4.98 PB quota on the largest and 4.38 PB provisioned on another, and the estate runs 38–40K job instances a day across ~2,500 scheduled jobs, on 3,400 Hadoop cores. Every job gets a dedicated VM that is destroyed on completion, so tenants are isolated and nothing idles. Platform state moves onto a zone-redundant managed MySQL — 32 vCores, 256 GiB — with its primary in one region and failover in another, replacing a single node that was backed up and patched by hand. I own the move off HashiCorp Vault and its retirement at the end of it. Secrets sync Vault → Key Vault throughout the transition, with a reconciliation pass that catches drift rather than trusting the sync — but the sync is scaffolding, not the destination. Consumers move class by class and Vault comes out when the last one is off it. Consumers were split by how they actually read secrets: direct API callers got a Key Vault client layer preserving the old logical paths and read/write/patch/delete semantics so they migrated with almost no code change, while cluster workloads kept their native Secret contracts — mounted as ephemeral files through the CSI driver, or synced back into real Kubernetes Secrets where an app still needed one, under their existing Secret names and keys — a data-governance frontend, an alerting and validation service, a capacity monitor, a log-retention job and the batch compute tier all kept working untouched. Pod identities bind to Entra workload identity, so nothing handles a credential. Retiring the old store is the risky half, so it runs on rules rather than nerve: one system is the source of truth at any moment and synchronisation is one-way, so a stale write cannot travel backwards into the store being drained. Staged migration modes let one consumer class move without the rest following, parity checks compare both stores before anything is cut, and the known failure modes are named and guarded — stale reads, incomplete listings, two patches racing the same secret — because each of those quietly produces a service that authenticates today and not tomorrow. Rollback stays available until the last consumer is off. Batch runtime came down from 21 to 11 minutes without breaking the six-hour report SLA. The sync started life as a weekend prototype — a Vault socket audit device piped through a TCP bridge into a RabbitMQ topic exchange, with a reconnecting publisher and bounded buffering, carrying event context only so the message bus never became a second place secrets live.

Alexa in-app purchase & subscriptions

Purchase, entitlement validation, billing, fulfilment. The design work is the failure path: a billing timeout drops into a retry queue and replays, and because entitlement is idempotent the customer is never double-charged on replay.

Hey Disney — partner integration

Backend design and launch readiness for partner purchasing, including the account-linking handshake between Alexa and the partner. Linking load time cut by roughly 85%.

Hey Disney! launch — About Amazon

Multi-room media & portability

Playback state synchronised across devices so media follows the listener. The trade-off is explicit: under a partition it favours availability and reconciles after, because stalled audio is a worse failure than a briefly stale position.

Multi-room audio in the Alexa app — The VergeConnecting Spotify and Apple Music — The Verge

Fleet Hub — cross-session agent observability

An always-on local dashboard that discovers fleets, sessions, agents, tasks and dependencies across Copilot, Claude and Codex session stores on both Windows and WSL, with no dependencies of its own — heterogeneous fleet JSON, agent trees and liveness signals normalised into one model with no central database behind it. New sessions register themselves through a session-start hook rather than by hand. A inuse.<pid>.lock file plus a live process check is what separates a running session from an abandoned one — so an ended fleet sorts away instead of sitting there claiming to be busy.

Broker-neutral messaging layer

Services needed to leave RabbitMQ for Azure Service Bus without every producer and consumer being rewritten around it — a nine-node broker estate spread across three regions, carrying fourteen production queues. The answer was a broker-neutral MessageBus contract in Scala with explicit Complete, delayed Retry and DeadLetter semantics — so unroutable publishes, delivery counts, bounded retries and dead-letter reasons are decided once instead of per service, while the legacy MQApi call paths keep compiling.

Alerting & validation modernization

Replacing a legacy alerting and validation system with open-source, Azure-aligned parts: a Go validation engine doing threshold and percent-difference logic, a Prometheus/Alertmanager stack with PromQL rules, an exporter and an OpsGenie mock. Alerting and validation were split into separate flows, and the whole thing runs on Docker and kind locally, wired to a shared pre-production environment — you can exercise a page before trusting it in production. Services either side are instrumented with OpenTelemetry metrics, traces and logs.

Where the hour actually went

Every execution carries eight timestamps, and end-to-end latency is measured from the hour the data was for — not from the moment the scheduler happened to pick it up. Measure it the usual way and a job delayed forty minutes in a queue still reports a fast run. Splitting the interval into scheduler, pending, queue and execution means a missed SLA points at the stage that caused it instead of starting an argument.

Capacity as a forecast, not an alarm

Every owned path reports space and object counts against quota, a seven- and thirty-day growth rate, and an estimate of how many days remain until 80, 90 and 100 percent. A threshold alarm tells you that you have a problem today; a growth curve tells you which team needs to be in a room three weeks from now. Re-evaluated every ten minutes, with the cache propagation delay treated as a known lag rather than a surprise.

Stack

Multi-agent orchestrationMCPCopilot extensibility TypeScriptPythonJavaGoScalaKotlinReact AzureAKSKey VaultTerraformCDKTFBicep AWSLambdaDynamoDBSQS/SNS KubernetesDockerSparkHadoop OpenTelemetryPrometheusService Bus

Education

Certificates