I’m obsessed with building the best agents, extracting every ounce of accuracy under realistic latency and token budgets, backed by extensive trace evals.
AboutTurning research on efficient agents into products people can rely on.
Experience
Hover a row for the detail.Click a row for the detail.
Engineering
Résumé ↓- 01Brainbox AutomationsAI EngineerDec 2025 – Now
- Deployed a safety-guardrailed AI coaching backend on GCP Cloud Run (FastAPI, Vertex AI) serving 300+ minor and adult athletes
- Cut CI/CD deploys from ~2 min to 10 s with GitHub Actions, keeping an SSH fallback
- Built SQL-aware agents over 50+ Supabase tables, answering 50% faster
- Shipped chat and VAPI voice AI for Earlibird (AU) handling 200+ chats a day and 5,000+ calls, with 80% resolved by AI and 21% booking conversion
- Summarised hour-long sales calls with a parallel map-reduce over Gemini on AWS Lambda
- 02Epistemy UKSoftware Engineer (AI) & Team LeadSep – Dec 2025
- Led the end-to-end build of an AI tutoring platform with a Nest.js backend (20+ endpoints) and a Next.js frontend
- Orchestrated multi-agent workflows on an event-driven Redis/BullMQ queue with fault tolerance
- Held 90%+ unit-test coverage with CI and pre-commit hooks, leading a team of 2
- 03CogniMind AIMachine Learning InternFeb – Apr 2025
- Raised VLM extraction accuracy by 20% through prompt engineering
- Sped up retrieval and inference by 10% with quantization and HNSW
- Automated MLOps with Dockerized Airflow and 5+ DAGs, cutting manual work by 60%
- Built Docker and GitHub Actions CI/CD that cut deploy time by 30%
- 04RapidsAI Machine Learning InternSep – Dec 2024
- Built a RAG chatbot (Streamlit, FastAPI) with contextual session management
- Reduced errors by 50% with chain-of-thought prompting
- Halved OpenAI API costs (−50%) with a complexity-based multi-model router
Research
Academic CV ↓- 01IPT Lab, NUSTResearcherJun 2026 – Now
- Leading 2 studies on redundancy and efficient inference in foundation models
- Showed similarity-based block removal (CKA, Procrustes) mostly reproduces an untrained network’s ranking (ρ ≥ 0.88)
- Certified early exit with learn-then-test and LoRA (7.3% of weights), saving 14–39% of encoder compute at 95% confidence
- Showed the guarantee breaks when SNR is estimated, and restored it by certifying on estimated groups
- 02Bradbury LabResearch Intern · remoteApr 2025 – Jan 2026
- Proposed a training-free layer-merging method based on Tucker decomposition to cut parameter count
- Analysed self-attention to test aligning Query and Key projections in efficient-by-design architectures
- Reviewed Transformer topology and parameter-efficient fine-tuning, focusing on weight sharing
- 03MachVis LabUndergraduate Research InternOct – Dec 2025
- Built an LLM-based factual-verification framework for the clinical accuracy of generated pathology reports, beyond BLEU/ROUGE
- Curated a challenge set of 100+ gigapixel whole-slide images with real artifacts and stain variation
- 04NUST · remoteUndergraduate Research InternJun – Sep 2025
- Combined MedSAM with meta-learning for few-shot dental radiograph segmentation, +12% on scarce disease classes
Inside the research
Measuring structural redundancy against proper controls, compressing models, and attaching statistical guarantees to compression decisions. Open to exploring agentic research on resource-constrained devices.
Blocks of similar layers look removable. Similarity predicts which ones no better than an untrained network.
Each band is an input through the layers, stopping where it is confident enough, with error ≤ α at 95% confidence.
A large teacher distilled through a narrow channel into a small task-specific student.
MedGemma 1.5 4B reads a slide into a fluent report, but gets treatment-critical facts wrong 41–67% of the time, and ROUGE-L can’t tell.
- Encoder compute saved by certified early exit
- 14–39%
- Spearman’s rho between similarity-based removal order and an untrained network’s
- ≥0.88
- Of the DeepSeek-VL2 (MoE) teacher’s performance retained by LiteDoc, on average
- 89.6%
- C1DOI ↗ (opens in a new tab)
LiteDoc: Distilling Large Document Models into Efficient Task-Specific Encoders
Tayyab, Taha, Adrian, Ulrich, Momina, Faisal, ICDAR 2026, Springer LNCS
- M1In prep.
Similarity Is Not Importance: On Measuring Representational Redundancy in Wireless Foundation Models
Taha, Draft on request
- M2In prep.
Certified Early Exit in a Wireless Foundation Model When the Signal-to-Noise Ratio Must Be Estimated
Taha, Draft on request
- M3Submitted
Accuracy of Craniometric Features in Gender Estimation Using Machine Learning Algorithms on University of Tennessee (UT) and Howells Datasets
Nuzhat, Taha, et al., 2026
Selected projects
- { 01 / 04 }
create-rag-app
GitHub ↗ (opens in a new tab)A Python CLI that scaffolds production-ready, Dockerized RAG apps from 13 Jinja2 templates, with 3 retrieval strategies and 8+ integrations.
- { 02 / 04 }
Bare-metal Prospector
GitHub ↗ (opens in a new tab)A framework-free SDR agent on a raw Python ReAct loop, with a syntax-directed parser that heals tool hallucinations and a three-tier memory.
- { 03 / 04 }
GPT-2 from scratch
GitHub ↗ (opens in a new tab)Pretrained GPT-2 from scratch in PyTorch, sped up inference with KV caching and speculative decoding, and instruction-tuned it on Alpaca.
- { 04 / 04 }
Agentic RAG
GitHub ↗ (opens in a new tab)A LangGraph planner and re-planner agent for retrieval, evaluated with RAGAS at over 95% faithfulness.
