Skip to content
Portfolio — 2026
Abhishek Tuteja

Abhishek
Tuteja

Open to Work

AI Engineer · LLM Systems

I build LLM applications, evals first. Ex-Deloitte, where I owned daily risk numbers behind Fortune 100 CTRM.

About

I'm an AI engineer building LLM applications: multi-agent systems and RAG, with the evals that decide whether they ship. Most of my work is on the parts that don't demo well: adversarial inputs, and the scores that quietly drift after a prompt change. The problems I care about live in what a model does when the easy path runs out.

That instinct came from Deloitte, where for two years I owned the daily risk numbers behind Fortune 100 commodity trading. Thousands of positions reconciled every morning, and a single break was mine to find and explain. When accuracy is audited every day, you learn to build for the failure cases first.

A Computer Science master's from Northeastern (3.96 GPA) sits under the engineering. What I've built with it is below.

AI & LLM

Multi-Agent SystemsLLM OrchestrationRAGPrompt EngineeringLLM-as-JudgeLLM EvaluationPrompt-Injection Defense

Languages

PythonTypeScriptSQLJava

Full-Stack & Data

Next.jsReactNode.jsSupabasePostgreSQLpgvectorPinecone

Testing & CI/CD

JestPlaywrightpytestGitHub ActionsDockerGit

Enterprise (SAP)

SAP S/4HANASAP ACMSAP BWSAP SignavioSOX / RCMUAT

Experience

01Jan 2025 - Apr 2026

Teaching Assistant

Northeastern UniversityOakland, CA

  • Teaching assistant for three core Computer Science courses: Introduction to Programming (CS 2000, Python and Pyret), Object-Oriented Design in Java (CS 3100), and Database Design (CS 3200).
  • Led graded code-walk reviews, assessed assignments, and held office hours, coaching students through programming, object-oriented design, and relational modeling.
02May 2022 - July 2024

Advisory Analyst

DeloitteBengaluru, IN

  • Owned the daily Stock Mark-to-Market risk report for a Fortune 100 agricultural commodities client, reconciling multi-commodity positions across SAP ACM and BW and validating long and short P&L and landed cost over thousands of records every day.
  • Drove daily open issues to near zero, investigating transaction breaks across the contract, load-capture, and settlement chain, reproducing defects in DEV, scoping fixes for ABAP developers, and validating through UAT.
  • Led end-of-day reporting for a three-analyst Center of Excellence and rolled out SAP ACM across LATAM, EMEA, and APAC, handling configuration localization and knowledge transfer. Deloitte Spot Award recipient.
  • Ran current- and future-state process workshops for a US oil and gas client on a Signavio ERP transformation, mapping process risks to controls and owning the SOX and Non-SOX Risk and Control Matrix (RCM) through stakeholder sign-off.
03Oct 2020 - Nov 2020

Machine Learning Intern

Steinn LabsRemote

  • Built an OCR pipeline in Python and TensorFlow that pulled structured data from scanned identity documents for database integration.
  • Lifted extraction accuracy by roughly 15% through preprocessing, model tuning, and validation against labeled samples.

Education

Master of Science in Computer Science

Sept 2024 - May 2026

Northeastern UniversityOakland, CA

GPA: 3.96

Algorithms, Database Management Systems, Programming Design Paradigm, Foundations of AI, Foundations of Software Engineering, Machine Learning, Generative AI, AI-Assisted Development

Bachelor of Technology in Mechanical Engineering (Minor: Data Science)

July 2018 - July 2022

Manipal Institute of TechnologyManipal, KA

GPA: 3.71

Projects

What I've been building.

01

PROVA

Upload model risk documentation, get an SR 11-7 compliance score and gap analysis in minutes.

Three concurrent LLM agents assess model documentation across 20 SR 11-7 elements, feeding a judge and orchestrator layer that runs cross-agent consistency checks, confidence-scored retries, and scoped re-assessment of disputed findings. The pipeline is hardened against adversarial input with prompt-injection defenses and a contrarian judge stance that counters self-enhancement and position bias. It ships on an evaluation harness of Jest and Playwright tests plus a custom score-drift eval that fails CI when a prompt change moves scores more than ten points. Output is a weighted compliance score with per-element gap analysis, evidence citations, and a downloadable PDF report.

Multi-AgentLLM EvalAnthropic APINext.js
02

CAPTRACK

Trade and portfolio analytics that get P&L right under messy, real trading conditions.

Ingests raw broker trade data, normalizes it into a consistent schema, and constructs positions with accurate realized and unrealized P&L. Built for the cases that break naive trackers: partial fills, multiple trades per asset, short positions, and multi-currency positions priced through real-time FX rates into a single base currency. A P&L accuracy testing framework guards calculation correctness, because the priority is deterministic, auditable accounting over flashy visualization.

Next.jsTypeScriptSupabasePostgreSQL

My Two Cents

Notes on building, breaking, and figuring things out.

01
6 min read

Why Our AI Agents Kept Lying About Their Scores

We built an app where three AI agents assess banking documents for regulatory compliance. Then we discovered they couldn’t do basic arithmetic. This is the…

model-validationmulti-agent-systems
Read article →
02
6 min read

Most Habit Apps Punish You for Missing a Day. Forge Doesn’t.

Instead of punishing missed days, Forge helps you understand your patterns, design momentum, and build identity through habits.Most habit apps make one…

technologyself-helphabitsatomic-habit
Read article →

Get in Touch

Have a question or want to work together?