Agentic security & governance
Scanning, RL driven red-teaming, and indexing composite agentic systems (LLM agents, MCP servers, tools, and skill files) before they ship, with findings mapped to Agentic Security Taxonomies.
Senior Applied Scientist & Tech Lead at Amazon AWS. I build the infrastructure that makes
Work at the intersection of large language models, reinforcement learning, and representation learning. Research into production across Amazon Bedrock and Amazon AgentCore, from first prototype to general availability.
Leading science for agentic AI science primitives: agent security and RL driven offensive red-teaming, agent customization and optimization, agentic search, and agent evaluations across task success, tool-use correctness, and safety adherence.
Talking with researchers, founders and others at frontier labs and agentic-AI startups. Especially interested in alignment & customization for LLMs and agents, security and safety for agentic systems, RL for tool use, and evaluation methodology.
My work sits at the platform layer: the primitives that let builders compose, secure, customize, search over, and evaluate populations of LLM agents in production.
Scanning, RL driven red-teaming, and indexing composite agentic systems (LLM agents, MCP servers, tools, and skill files) before they ship, with findings mapped to Agentic Security Taxonomies.
Mining OpenTelemetry traces from live agent runs to auto-generate skill files, tighten system prompts and tool descriptions, and train reinforcement-learning policies for optimization and customization.
Intelligent model routing with calibrated quality predictors and user-controlled cost/latency tradeoffs; automatic prompt optimization; trajectory evaluation and safety adherence for composite agents.
Work on Invariant and equivariant architectures for baking in discrete and continuous symmetries into neural networks: Janossy Pooling, Relational Pooling, equivalence of positional and structural embeddings, and ESAN subgraph GNNs.
Selected products I led or contributed to at AWS, from first prototype to general availability, across Amazon AgentCore, Bedrock, and DataZone.
Platform primitive for automatically customizing and optimizing deployed AI agents. Mines OpenTelemetry traces from live runs to auto-generate skill files, tighten system prompts and tool descriptions, and train RL policies, turning raw telemetry into continuous improvements.
Managed registry for AI agents: discover, govern, and secure agents and MCP servers across the enterprise. Includes agent search, policy controls, and pre-deployment security scans for composite agentic systems.
Production evaluation framework for AI agents covering task success, tool-use correctness, trajectory quality, and safety adherence. Continuous monitoring against reliability and policy gates before and after deployment.
Routes each request to the most suitable model within a model family. Calibrated quality predictor, user-controlled quality-cost tolerance, offline and online eval harness. Scaled to 1M+ requests/week with 40%+ cost reduction vs. top model at <150 ms routing latency.
Automatically rewrites prompts for better performance across Claude, Llama, Nova, Mistral, Titan, and DeepSeek. Contributed to rewrite strategies, per-task quality evaluation, and task coverage spanning summarization, RAG/QA, classification, function calling, and reasoning.
LLM-based documentation for enterprise data catalogs. Generates summaries and column descriptions without access to column contents. Schema-aware routing, iterative refinement, and LLM-as-judge evaluation across 1,000+ column tables and 25+ enterprise domains.
Generative AI layer for Amazon DataZone that proposes business-friendly names and descriptions for technical assets, so analysts can discover data by what it means, not by cryptic column names, at enterprise scale.
Click any paper to read the abstract. Full list of papers and patents on Google Scholar.
Casts workflow generation as Bayesian inference over a posterior distribution on workflows: a sampling framework that builds workflows step-by-step with parallel look-ahead rollouts and a sequential in-loop refiner. Improves accuracy by up to 9 points over state-of-the-art and 65 points over zero-shot prompting.
Multi-task fine-tuning approach for related table discovery in data lakes, enabling LLMs to jointly reason over schema, content, and provenance signals to find semantically related tables at enterprise scale.
Advances LLMs as open-domain table reasoners by combining retrieval-augmented evidence gathering with iterative chain-of-thought reasoning over tables, enabling strong performance on complex table QA tasks without task-specific fine-tuning.
Expresses permutation-invariant functions of sequences as the average of a permutation-sensitive function over all reorderings, unlocking the full literature of RNNs and CNNs for invariant tasks. Introduces tractable approximations via canonical orderings, k-ary interactions, and stochastic optimization.
Whether you’re a researcher, founder, or anyone else: if you’re working on hard problems at the frontier of agentic AI, I’d love to hear from you.