Multi-agent orchestration · Microsoft AI, productized independently
One agent is an assistant. A fleet is a force multiplier.
Ran a real Azure migration with up to eight agents in parallel — minutes, not a night — while the human only reviewed.
Watch the workflow — 1:48one sequential agent versus a fleet, on a real Azure migration: dependency classification, eager fan-out, warm reuse, and up to eight agents in parallel while the human only reviews.A task decomposes, fans out to agents, and those agents fan out again — nesting as deep as the work needs. Reconciliation is dynamic and runs across fans, not just up the tree: sibling branches settle with each other as results land, then fold level by level, one branch owning any shared resource, until a single result passes the verification gate. Shipped on GitHub Copilot, Copilot CLI, and VS Code from one canonical core.
Inside the runtimeInside: the scheduler decomposes and right-sizes each task, a lock table guarantees only one agent owns a shared resource at a time, and agents are claimed from a warm pool under a lease — reused rather than cold-started, capped so idle capacity cannot pile up, and reaped on TTL when a lease goes stale. The reconciler merges what returns, and nothing reaches "complete" without passing the gate. Work is modelled explicitly — tasks, dependencies, owners, resource claims, status and verification evidence — so a run in flight stays inspectable and recoverable rather than opaque, and partial sub-agent output can never be reported as finished work.
AI systems I've shipped
Armada Multi-agent orchestration — nested fan-out where agents recursively decompose further, across-fan reconciliation where sibling branches settle with each other rather than only folding up the tree, single-owner resource scoping, and verification gates before completion.
Three surfaces, one core VS Code extension (@armada chat participant, fleet tree, agent/task metadata, Fleet Hub dashboard) and the Copilot CLI plugin (/armada, armada-fan-out skill, MCP fleet observability) — generated from a canonical core with surface adapters so the integrations cannot drift.
/ground-rule Persistent AI operating standards that survive across Copilot sessions. The interesting part was getting there: rather than patching the CLI bundle or leaning on an unsupported custom-command directory, I reverse-engineered Copilot's local marketplace plugin schema and built on the supported path. A canonical JSON store syncs into Copilot instructions through protected markers, with atomic writes, retry/backoff and reinstall-safe migration so a re-install cannot silently drop your rules.
/ado-pr Governed Azure DevOps pull requests — list, inspect, approve, and merge by URL or ID, with team/author/repo/status filtering and structured output for automation. Two-phase approval: read-only inspection, explicit confirmation, short-lived action tokens, policy checks, guarded bypass.
Fleet Hub An always-on local dashboard for every fleet, session, agent and task — discovered across Copilot, Claude and Codex session stores on Windows and WSL with no dependencies, bootstrapped by a session-start hook, and kept honest by a pid-lock liveness check so an ended session never reports as running.
Agent economics & privacy Lease-guarded claims from a warm pool, capped idle capacity, TTL reaping, dependency-aware scheduling, retry boundaries and model right-sizing; local-first fleet state with sanitized summaries, path redaction, and data minimisation of Copilot session artifacts.
Platform toolkit 28 Terraform / CDKTF skills, each turning a repeatable platform chore into a governed self-service workflow — built on Terraform, CDKTF, Bicep, Azure DevOps and the Azure CLI, against AKS, ACR, Key Vault, Workload Identity and the Secrets Store CSI Driver, in TypeScript, Node.js and PowerShell.
Armada Fan-Out
Terraform Security Baseline Applier
State Drift Resolver
Plan Reviewer
Terraform PR Quality Evaluator
CDKTF PR Quality Evaluator
State Migrator
State Backend Configurator
Terraform State Storage NSP Applier
Terraform Provider Version Policy
Terraform Provider Upgrade
Terraform-to-CDKTF Converter
Terraform-to-Modules Migrator
CDKTF Module Binding Generator
CDKTF Project Bootstrapper
New Module Scaffolder
Module Usage Generator
Module Version Pinning
Module Catalog Lookup
Cost Impact Estimator
Environment Promotion Scaffolder
Naming Convention Applier
Tagging Standard Applier
Azure VMSS Org Standards
ADO Service Connection Creator
Providers JSON Generator
Terraform Error Troubleshooter
Skill Test Harness Scaffolder
Experience
MicrosoftJun 2026 — now
Senior Software Engineer, Microsoft AI. Joined an in-flight on-prem → Azure program; what I have owned since: Armada, Fleet Hub, the Azure networking layer and Private Link, Vault → Key Vault with drift reconciliation, and AKS job execution on ephemeral per-job VMs.
Amazon2019 — 2026
Software Development Engineer. Tier-1 push notifications at ~6M events/min, 99.99%+ availability. Alexa in-app purchase 0 → ~$4M/yr. Hey Disney integration. Add to Delivery notifications for Prime members. ~$14M/yr cut from the Rekognition fleet.
Built the team's client-side metrics and alarming from zero — replacing hand-maintained monitors with generated, version-controlled ones — cut cross-service call volume by half, contributed components to a shared library adopted by 10+ teams, and led Prime Day and Q4 peak readiness. Became the only certified security reviewer in the org while still at SDE1.
OOracle2017 — 2019
Senior Software Engineer. Java/Spring Boot microservices on the Advanced Support Platform; production errors down ~45% via JVM profiling and query tuning. Technical lead for 10+ engineers. Product docs
Senior Software Engineer. Artik marketplace features and partner portals — ~27% revenue growth.
A&MTexas A&M–Kingsville2014 — 2015
Research Assistant on the Software Visualization project — a web IDE rendering static structure as UML and dynamic behaviour from runtime stack traces in D3 — and Graduate Teaching Assistant for Operating Systems and Mobile Application Development, where I built the automated attendance and grading system the courses ran on.
LLocuz Enterprise Solutions2013 — 2014
Software/Cloud Engineer, Innovative Development Team. A cross-vendor cloud management API — configuration, installation, reporting and analysis — validated in EMC² VSPEX and NetApp FlexPod labs; a cluster manager with job submission and tracking that worked independently of the underlying cluster suite; and aCube, centralised licence management over a FlexLM server.
IIT Bombay2011 — 2013
Virtual Labs, Chemical Engineering. Built the Healthcare Research Consortium platform and the pre/post-viva modules, ran the lab's NAS and Linux compute server, and taught FOSS workshops for the Spoken Tutorial project under MHRD's National Mission on Education through ICT.
Beyond the codebase
Where the work reaches people who do not report to me.
12.5KFollowing on LinkedInwriting about agent systems, orchestration and what actually breaks in production
O'ReillyAI Agents in Pythonpublished technical work on building agent systems
40+Workshops taughtLinux, PHP/MySQL and Python, run for IIT Bombay's Spoken Tutorial project under the Government of India's National Mission on Education through ICT
6 yrsAmazon Security Certifierdesign-review authority to block another team's production launch on security and readiness
Leading
Authority, standards and the times somebody had to pick it up.
Authority to stop another team's launch
Amazon Security Certifier from 2020 — and the only certified security reviewer in the org while still at SDE1. The role carries design-review authority over other teams' production readiness, which means the power to say a launch is not going out. Exercised across multiple services, more than the bar required.
Picked it up when the team halved
After roughly 60% of the team left in one quarter, I took Prime Day and Q4 peak readiness onto myself rather than waiting for it to be assigned, and carried the component work the wider org was depending on. Nobody appointed me to either.
Changed the plan by writing the argument down
Proposed splitting a service the team had not planned to split, wrote the design documents for both halves in parallel, and took them through review with the data behind the trade-offs. My manager was not initially convinced. The design went in as I argued it, and the second service exists because of that document.
Standards other teams pick up
Reusable service and infrastructure templates adopted across teams; components contributed to a shared library used by 10+ teams; a 28-skill platform toolkit that turns repeatable chores into governed workflows; and a persistent engineering-standards system so a team's operating rules survive between sessions rather than living in someone's head.
Engineers, directly
Technical lead for a 10+ engineer team at Oracle — design reviews, code reviews, performance debugging. Mentoring and cross-team design reviews at Amazon. At Microsoft I own each team's move onto the platform end to end: discovery, scoping, design review, hands-on migration of their workloads, and cutover support.
Teaching, before any of it
40+ open-source workshops taught across colleges in India for IIT Bombay's Spoken Tutorial project under a Government of India programme, with standing authorization to run them anywhere in the country — and a Graduate Teaching Assistant post at Texas A&M–Kingsville for Operating Systems and Mobile Application Development.
Also built
Beampersonal
My one-year-old lost the TV remote, so I rebuilt the TV instead. Overnight, on a discontinued Meta Portal TV stripped of every app — no docs, no root, no store. By morning: a from-scratch launcher running 25+ streaming apps natively (Netflix, Prime Video, Disney+, Hulu, Spotify), 700+ live channels in 11 languages, and a content-first UI with auto-playing hero trailers and Continue Watching. Daily driver ever since.
Paste a recruiter thread — or connect the mailbox — and it extracts company, role, stage and contact into a tracked record, then reads the ICS attachment or the Outlook calendar event over Microsoft Graph and books the interview itself. There is a deterministic heuristic fallback for when no model key is configured, so the app degrades instead of failing.
Encryption is per-field, not per-database: notes, salary, recruiter identity, résumé text, raw email bodies, summaries and sentiment each AES-256-GCM encrypted at rest, IMAP passwords and OAuth tokens under a separate salt — and encrypt-on-write / decrypt-on-read live in the storage layer, so no call site can forget to do it. React and Express over Drizzle and PostgreSQL, with a 60-page high-level design behind it.
Scorebookpersonal
Offline-first cricket scoring: a scorer at a ground with no signal records a full match in airplane mode, and it reconciles on reconnect — the hard part being conflict and duplicate resolution, not the UI. Vanilla-JS PWA over a Node/SQLite server with a Capacitor iOS shell, a DLS rules engine, 26 test suites, and CI deploying to Azure. ~67K lines, 282 commits, all his own.
Siftpersonal
A macOS desktop app that files documents by what is inside them rather than what they are named — text pulled out of PDFs and Word files, Vision-framework OCR when a PDF is image-only, then classified, deduplicated by content hash with a keep-preference ranking, and balanced across local, iCloud, Google Drive and OneDrive accounts. The sync is deliberately timid: additive, newer-wins, and it never propagates a delete — the losing copy is kept as a .conflict file instead of disappearing. Electron shell over a Python engine with a Swift OCR helper, ~4.8K lines.
Job-hunt automationpersonal
A scheduled scanner screens job boards against a role rubric, a generator edits the master .docx in place so fonts and spacing stay byte-identical rather than being re-derived, and the submit flow always stops at a human-review gate before anything is sent. Built on the open-source career-ops pipeline (santifer/career-ops, MIT).
SquadUppersonal
A sports-club operations platform, built as a TypeScript monorepo — NestJS and Prisma over PostgreSQL, React/Vite on the web, Expo on mobile, with contracts shared across all three so the clients cannot drift from the API. Covers the unglamorous half of running a league: clubs, teams, players, grounds, umpires, match formats, public tournaments, scoring, and the parts most tools skip — committee elections, and promotion and relegation between divisions.
Kubernetes Hyperspacepersonal
A cluster you can fly through rather than grep. A local-first FastAPI backend shells out to kubectl and a React force-graph renders the topology in three dimensions — resource identity, health, dependency paths and logs, refreshing itself as the cluster moves. Drill-down, exec, scale, restart and a command palette are wired in, but the write actions are deliberately bounded, access stays local, and credentials are masked rather than shipped anywhere.
Code Universepersonal
Point it at a local directory and it compiles an architecture map you can walk — scanning incrementally, detecting and consolidating projects across a multi-project workspace, then giving you breadcrumbs, navigation history and jumps along real dependency edges. Raw code stays local by default and is never executed, and rendering is bounded on purpose so a large workspace explores rather than hanging the browser.
Copilot CLI Cockpitpersonal
A live cockpit for CLI sessions where zooming the active pane does not shove the navigation and command chrome around it. The trick was moving zoom into the DOM and holding the terminal chrome static, plus an embed mode that renders the stage alone — driven by a control HTTP channel and a dock host wrapped around per-connection child processes.
Architecture I designed
Systems I owned end to end — drawn as they actually behave, including the failure paths.
Tier-1 push notification platform
The platform behind notifications like Add to Delivery, which tells a Prime member their order can still be changed — so it is only useful if it arrives before the van does. Producers fan into SNS, SQS absorbs the spike as back-pressure, Fargate workers drain it, and every write is idempotent so a replay cannot double-send. What fails retries; what keeps failing lands in the DLQ instead of blocking the queue.
The multi-region Hadoop and Spark platform where all bidding data lands, reached over Private Link — sitting on a dependency map of batch compute, an LDAP auth bridge, a MySQL metadata store, object storage, message brokers, Redis, ZooKeeper and the operational log pipelines — where single tenants run to petabytes, 1.05 PB live against a 4.98 PB quota on the largest and 4.38 PB provisioned on another, and the estate runs 38–40K job instances a day across ~2,500 scheduled jobs, on 3,400 Hadoop cores. Every job gets a dedicated VM that is destroyed on completion, so tenants are isolated and nothing idles. Platform state moves onto a zone-redundant managed MySQL — 32 vCores, 256 GiB — with its primary in one region and failover in another, replacing a single node that was backed up and patched by hand. I own the move off HashiCorp Vault and its retirement at the end of it. Secrets sync Vault → Key Vault throughout the transition, with a reconciliation pass that catches drift rather than trusting the sync — but the sync is scaffolding, not the destination. Consumers move class by class and Vault comes out when the last one is off it. Consumers were split by how they actually read secrets: direct API callers got a Key Vault client layer preserving the old logical paths and read/write/patch/delete semantics so they migrated with almost no code change, while cluster workloads kept their native Secret contracts — mounted as ephemeral files through the CSI driver, or synced back into real Kubernetes Secrets where an app still needed one, under their existing Secret names and keys — a data-governance frontend, an alerting and validation service, a capacity monitor, a log-retention job and the batch compute tier all kept working untouched. Pod identities bind to Entra workload identity, so nothing handles a credential. Retiring the old store is the risky half, so it runs on rules rather than nerve: one system is the source of truth at any moment and synchronisation is one-way, so a stale write cannot travel backwards into the store being drained. Staged migration modes let one consumer class move without the rest following, parity checks compare both stores before anything is cut, and the known failure modes are named and guarded — stale reads, incomplete listings, two patches racing the same secret — because each of those quietly produces a service that authenticates today and not tomorrow. Rollback stays available until the last consumer is off. Batch runtime came down from 21 to 11 minutes without breaking the six-hour report SLA. The sync started life as a weekend prototype — a Vault socket audit device piped through a TCP bridge into a RabbitMQ topic exchange, with a reconnecting publisher and bounded buffering, carrying event context only so the message bus never became a second place secrets live.
Alexa in-app purchase & subscriptions
Purchase, entitlement validation, billing, fulfilment. The design work is the failure path: a billing timeout drops into a retry queue and replays, and because entitlement is idempotent the customer is never double-charged on replay.
Hey Disney — partner integration
Backend design and launch readiness for partner purchasing, including the account-linking handshake between Alexa and the partner. Linking load time cut by roughly 85%.
Playback state synchronised across devices so media follows the listener. The trade-off is explicit: under a partition it favours availability and reconciles after, because stalled audio is a worse failure than a briefly stale position.
An always-on local dashboard that discovers fleets, sessions, agents, tasks and dependencies across Copilot, Claude and Codex session stores on both Windows and WSL, with no dependencies of its own — heterogeneous fleet JSON, agent trees and liveness signals normalised into one model with no central database behind it. New sessions register themselves through a session-start hook rather than by hand. A inuse.<pid>.lock file plus a live process check is what separates a running session from an abandoned one — so an ended fleet sorts away instead of sitting there claiming to be busy.
Broker-neutral messaging layer
Services needed to leave RabbitMQ for Azure Service Bus without every producer and consumer being rewritten around it — a nine-node broker estate spread across three regions, carrying fourteen production queues. The answer was a broker-neutral MessageBus contract in Scala with explicit Complete, delayed Retry and DeadLetter semantics — so unroutable publishes, delivery counts, bounded retries and dead-letter reasons are decided once instead of per service, while the legacy MQApi call paths keep compiling.
Alerting & validation modernization
Replacing a legacy alerting and validation system with open-source, Azure-aligned parts: a Go validation engine doing threshold and percent-difference logic, a Prometheus/Alertmanager stack with PromQL rules, an exporter and an OpsGenie mock. Alerting and validation were split into separate flows, and the whole thing runs on Docker and kind locally, wired to a shared pre-production environment — you can exercise a page before trusting it in production. Services either side are instrumented with OpenTelemetry metrics, traces and logs.
Where the hour actually went
Every execution carries eight timestamps, and end-to-end latency is measured from the hour the data was for — not from the moment the scheduler happened to pick it up. Measure it the usual way and a job delayed forty minutes in a queue still reports a fast run. Splitting the interval into scheduler, pending, queue and execution means a missed SLA points at the stage that caused it instead of starting an argument.
Capacity as a forecast, not an alarm
Every owned path reports space and object counts against quota, a seven- and thirty-day growth rate, and an estimate of how many days remain until 80, 90 and 100 percent. A threshold alarm tells you that you have a problem today; a growth curve tells you which team needs to be in a room three weeks from now. Re-evaluated every ten minutes, with the cache propagation delay treated as a known lag rather than a surprise.
Stack
Multi-agent orchestrationMCPCopilot extensibilityTypeScriptPythonJavaGoScalaKotlinReactAzureAKSKey VaultTerraformCDKTFBicepAWSLambdaDynamoDBSQS/SNSKubernetesDockerSparkHadoopOpenTelemetryPrometheusService Bus
Education
M.S. Computer ScienceTexas A&M University–KingsvilleConferred May 2016 · compilers, distributed systems, cryptography, cloud computing
B.Tech Computer Science & EngineeringRajasthan Technical University — Arya College of Engineering & IT, Jaipur2012 · First Division
Certificates
Red Hat Certified Engineer (RHCE)Red Hat · RHEL 5 · Oct 2010
Security: Network Services — Certificate of ExpertiseRed Hat · EX333 · Dec 2011
Directory Services & Authentication — Certificate of ExpertiseRed Hat · EX423 · Jul 2011
DEV 360 — Apache Spark EssentialsMapR Academy · Feb 2017
DEV 361 — Build & Monitor Apache Spark ApplicationsMapR Academy · Feb 2017
DEV 362 — Create Data Pipelines Using Apache SparkMapR Academy · Feb 2017
DEV 301 — Developing Hadoop ApplicationsMapR · Jul 2015
HDE 100 — Hadoop EssentialsMapR · Feb 2015
Architecting with Google Kubernetes Engine — WorkloadsGoogle Cloud · May 2026 · verify