Senior Software Engineer / Applied AI / Backend & full-stack systems
PrateekMulye
The riverwalk, Chicago
I build AI applications on 11+ years in software engineering, most of it backend and full-stack systems. At work I built production systems that ran on live traffic, in Java and Elixir. In my own projects I build the parts around the model: streaming, retrieval, routing and evaluation.
Personal projects, live online. Each one leaves the decision with a person and says where the model falls short.
Banking operations
Payment Exception Desk
Matches incoming credits to expected payments. Exact amount, currency and date checks narrow the candidates first; a local MiniLM model then ranks descriptions for an analyst to confirm.
LimitNot evaluated on a permissioned payment dataset. Every match needs an analyst decision.
Compares a building’s interval electricity readings with its operating schedule, and tests a learned model against simple calendar baselines before flagging anything.
LimitOn the three real building files tested, the model did not beat the baselines, so the tool withheld model flags.
Runs Whisper in the browser to find where a caption track disagrees with the recorded speech, then puts each suggestion in a cue editor beside the recording.
LimitAudio comparison needs PCM WAV input. Accuracy and time savings are not measured.
Assay · CP044 community project · personal AI contribution
Research you can inspect and compare.
In Assay, my contribution to the CP044 community project, I built Python and LangGraph workflows that bring analyst, bull, bear and risk perspectives into a streamed research report. The project also includes an evaluation harness that compares the debate workflow with a simpler one, including token use and estimated cost.
The decision and its limits
Decision. Make the extra stages earn their place: compare their outputs and cost with a simpler workflow before claiming they improve the answer.
Limits. A contribution to a community project, not sole authorship. Pinned source review, not a reproduced live-model run. Fixed judge order and excluded failed pairs limit how far the harness results can be read.
Illustrative. The pinned source shows the stages; no live run is claimed.
This page · System One router
The navigation on this page is a fine-tuned decision model.
Ask a question in the bar above, or paste a role into “Let’s talk”. Laya, an open-weights decision model (Apache-2.0), answers typed questions in one pass instead of writing text: which section fits, is this about hiring, is it off-topic. I fine-tuned it for this page on 1,655 synthetic questions and job descriptions written by a local model, and it runs in a private Hugging Face Space at no cost, answering only requests that carry this site’s keys. A confidence gate decides whether to go, offer two options or decline. If Laya is slow or down, a keyword classifier in your browser answers instead, and the trace says which one did. Every sentence it shows was written in advance and links to its source.
Measured on 60 held-out questions and 10 job descriptions
Keywords
Laya as released
Laya fine-tuned
Right section (top 1)
29/6036 to 61%
26/6032 to 56%
48/6068 to 88%
Jumped to the right section
29/6036 to 61%
10/609 to 28%
45/6063 to 84%
Jumped to a wrong section
5/604 to 18%
5/604 to 18%
6/605 to 20%
Fit brief: stated requirements
172/18092 to 98%
63/18028 to 42%
175/18094 to 99%
The test set was written by a separate agent that never saw any engine’s answers, in English and other EU languages, and frozen before the first run. Fine-tuned against keywords, paired exact McNemar test: top 1 p = 0.0003, right jumps p = 0.0015, wrong jumps p = 1. Before scoring I set a rule: ship only if wrong jumps were no more frequent than with keywords. At 6 against 5 it missed that rule by 1 case; I shipped it anyway and say so here. The fit brief asks Laya everything except work location, where it scored below the keywords; all 21 brief questions take 1.9 s on the Space. Ranges are 95% intervals.
How a question travels. The same path appears live in the trace when you ask.
Streaming inference · ChatFormula1 v2, in development
What happens when an event arrives twice?
An LLM answer reaches the user as a stream of events, and streams repeat, skip and arrive out of order. This console runs the original React client reducer from ChatFormula1 v2 on synthetic events in your browser. No model, API or database call is made.
The foundation · 11+ years in software engineering
Measured at work.
15M+
US and Canadian company records normalized into one canonical shape
HG Insights
30M+
PostgreSQL rows partitioned and indexed; dashboard latency down more than 60%
HG Insights
80K+
contracts through an idempotent Elixir and Oban Pro pipeline; reporting 60% faster
HG Insights
12K+
transactions per second from OpenResty service virtualization, in test environments
Capital One, via Cognizant
System design
Three systems, drawn out.
Diagrams are illustrative. They show the design, not production traces.
Capital One, via Cognizant · 2018–2021
When a message fails
In production, on live traffic, I owned the Kafka workflows for financial-transaction messages, built with Spring Boot and DynamoDB: dead-letter handling, bounded retries with backoff, and idempotent replay. Operators could recover a bad message without processing a transaction twice.
I also helped split a legacy Java application into contract-first services that teams could release on their own, and built the OpenResty layer that stood in for unavailable downstream systems during performance tests.
The decision and its limits
Decision. Separate retryable failures from work that needs inspection. In an AI application I would assess every tool action before making it retryable, because a repeated request can have a real external effect.
Limits. Career account. Idempotent replay is not a universal exactly-once delivery guarantee.
Java
Spring Boot
Kafka
DynamoDB
OpenResty
Bounded retry, then park and replay. Red marks the failure path.
HG Insights · 2021–2025
Tables that outgrew their indexes
I scaled PostgreSQL beyond 30 million rows with partitioning, composite indexes and read replicas. Customer-facing dashboard latency fell by more than 60%.
On the same team I built an idempotent contract-processing pipeline in Elixir and Phoenix with Oban Pro. It handled 80K+ records and cut reporting turnaround by 60%.
A production lesson I still use
A PostgreSQL database ran out of storage and took its application offline. I scaled it to restore service. The team traced the growth to retained Oban job records and added retention, scheduled cleanup and disk-usage alerts. That changed what I checked first in monitoring.
My account of the response and the team changes that followed. No production logs are published.
PostgreSQL
Elixir / Phoenix
Oban Pro
Read replicas
Partition the writes, route the reads.
HG Insights · 2021–2025
One shape for 15M+ records
I owned the architecture and phased rollout of a geo-standardization service with Data Solutions and downstream teams. It normalized 15M+ US and Canadian company records into canonical shapes used by ingestion and interface workflows.
A companion job system ran scheduled, on-demand and event-triggered company-spend refreshes across corporate hierarchies, moving delivery from a manual monthly batch to a weekly cadence.
The decision and its limits
Decision. Standardize at the ingestion boundary. I bring the same discipline to preparing data for retrieval, while evaluating retrieval quality separately.
Limits. Career account. Employer data and implementation are not public.
Data contracts
Elixir
PostgreSQL
Phased rollout
Illustrative inputs. One canonical output.
Professional record
Built over time.
At Agilent I built internal Gemba reporting and action-tracking apps with Power Apps and Power Automate, support manufacturing operations and recipe authoring, and help colleagues adopt AI.
Employment, most recent first
Dates
Employer
Role
City
2026 – now
Agilent Technologies
Senior Software Engineer, Manufacturing
Carpinteria, CA
2021 – 2025
HG Insights
Software Engineer
Santa Barbara, CA
2018 – 2021
Capital One, via Cognizant
Software Engineer
Vienna, VA
2018
Tuutkia
Software Engineer, part-time
San Jose, CA
2017 – 2018
Illinois Department of Public Health
Software Engineer Intern
Springfield, IL
2013 – 2015
AurionPro Solutions
Software Engineer
Pune, India
Where the work and study happened. Equal Earth projection; marks show city centres, not travel routes. Land: Natural Earth.
M.S. Computer ScienceUniversity of Illinois at Springfield · 2018
Post-Graduate Diploma, Advanced ComputingC-DAC, Pune · 2013
Generative AI with Large Language Models (non-credit course)DeepLearning.AI and AWS (Coursera) · Feb 2026 · verify ↗
I hike when I travel, and I will walk a city end to end before I take a car through it. Languages are a long-running project: English, Hindi and Marathi from home, some Italian, and German, A1 to A2, learning.