Report / 2026 / seven systems
Abdul Samad Zeeshan
Backend and data systems, some with a model inside
- projects
- 7
- demos, no login
- 7
- numbers reported
- 14
- reported as not good enough
- 2
Skills, linked to the project that proves each one.
AI / ML
- PythonWarden, Tarn, Proving, Parley, Triage
- LoRA fine-tuningTriage
- Knowledge distillationTriage
- Model evaluationTarn, Proving, Parley, Triage
- CalibrationTarn, Proving, Docket
- LangGraphWarden
- Model Context ProtocolWarden
- Tool-using agentsWarden, Tarn, Proving
- Schema-checked outputTriage, Docket
- Speech recognition and TTSParley
- Grounded repliesParley
- Amazon BedrockDocket
- Fraud and anomaly detectionTally, Tarn
Systems
Web / backend
Infrastructure
The rule
I finished my Computing Science degree at the University of Alberta in June 2026. I build backend services, data pipelines and systems with a language model somewhere inside. Every project below solves one real problem, and its page shows what actually happened when I tested it.
Projects / 01 to 07 / my top pick first
Tally
A small bank on a double-entry ledger that never loses or doubles a cent, even when the network fails mid-payment.
Transfers pushed through 7 runs on Kubernetes while the network and database were made to fail. Not one cent lost or doubled.11,054none lost
Warden
Guards changes to live software and cannot be talked round by a tricked AI assistant.
Attackers tried 84 ways to trick the AI into approving a bad change. None got through.0 of 84none got through
Tarn
Scores a billion real logins from a US national lab for signs of a known attacker, then an AI analyst sorts the alerts.
Days of real attacker activity caught when an analyst reads only 100 alerts a day. The first version caught 21. Most still got through.29 of 181most still got through
Proving
Tests each new version of an AI agent on thousands of made-up customers, then says ship or hold.
Hostile fake customers got the old Warden to make 1.3 unwanted server calls per conversation. The new one made 0.6. Verdict: ship.1.301 to 0.607ship
Parley
Answers the phone for a property agency in English and Gulf Arabic, books viewings, and only states facts the listings database returned.
Every time the booking system failed mid-call, all 232 times, Parley said so and offered a callback instead of inventing a price or address.232 of 232never invented a fact
Triage
A small offline model that sorts support emails, says how sure it is, and sends only the unsure ones to a bigger model.
Support emails given the right urgency, out of 1,200 it had never seen. The same model before training: 38.3%. Always guessing high: 42.4%.87.2%up from 38.3%
Docket
Reads receipt photos into checked data and says which receipts can skip a person at a 1 percent error budget.
Test receipts the system trusted enough to skip a person, at a 1% error budget. Of their 32 fields, none were wrong.11 of 149none wrong
Tools / 11 / used in the seven
PythonJavaTypeScriptSQL and PostgreSQLKubernetesAWS Lambda and CDKSpark and DuckDBLangGraph and MCPLoRA fine-tuningTLA+ and z3Evaluation and calibration






