I’m obsessed with building the best agents, extracting every ounce of accuracy under realistic latency and token budgets, backed by extensive trace evals.
AboutTurning research on efficient agents into products people can rely on.
Experience
Hover a row for the detail.
Engineering
Résumé ↓- 01Brainbox AutomationsAI EngineerDec 2025 – NowDeployed a safety-guardrailed AI coaching backend on GCP Cloud Run (FastAPI, Vertex AI) serving 300+ minor and adult athletes
- 02Epistemy UKSoftware Engineer (AI) & Team LeadSep – Dec 2025Led the end-to-end build of an AI tutoring platform with a Nest.js backend (20+ endpoints) and a Next.js frontend
- 03CogniMind AIMachine Learning InternFeb – Apr 2025Raised VLM extraction accuracy by 20% through prompt engineering
- 04RapidsAI Machine Learning InternSep – Dec 2024Built a RAG chatbot (Streamlit, FastAPI) with contextual session management
Research
Academic CV ↓- 01IPT Lab, NUSTResearcherJun 2026 – NowLeading 2 studies on redundancy and efficient inference in foundation models
- 02Bradbury LabResearch Intern · remoteApr 2025 – Jan 2026Proposed a training-free layer-merging method based on Tucker decomposition to cut parameter count
- 03MachVis LabUndergraduate Research InternOct – Dec 2025Built an LLM-based factual-verification framework for the clinical accuracy of generated pathology reports, beyond BLEU/ROUGE
- 04NUST · remoteUndergraduate Research InternJun – Sep 2025Combined MedSAM with meta-learning for few-shot dental radiograph segmentation, +12% on scarce disease classes
Inside the research
Measuring structural redundancy against proper controls, compressing models, and attaching statistical guarantees to compression decisions. Open to exploring agentic research on resource-constrained devices.
Blocks of similar layers look removable. Similarity predicts which ones no better than an untrained network.
Each band is an input through the layers, stopping where it is confident enough, with error ≤ α at 95% confidence.
A large teacher distilled through a narrow channel into a small task-specific student.
MedGemma 1.5 4B reads a slide into a fluent report, but gets treatment-critical facts wrong 41–67% of the time, and ROUGE-L can’t tell.
- Encoder compute saved by certified early exit
- 14–39%
- Spearman’s rho between similarity-based removal order and an untrained network’s
- ≥0.88
- Of the DeepSeek-VL2 (MoE) teacher’s performance retained by LiteDoc, on average
- 89.6%
- C1DOI ↗ (opens in a new tab)
LiteDoc: Distilling Large Document Models into Efficient Task-Specific Encoders
Tayyab, Taha, Adrian, Ulrich, Momina, Faisal, ICDAR 2026, Springer LNCS
- M1In prep.
Similarity Is Not Importance: On Measuring Representational Redundancy in Wireless Foundation Models
Taha, Draft on request
- M2In prep.
Certified Early Exit in a Wireless Foundation Model When the Signal-to-Noise Ratio Must Be Estimated
Taha, Draft on request
- M3Submitted
Accuracy of Craniometric Features in Gender Estimation Using Machine Learning Algorithms on University of Tennessee (UT) and Howells Datasets
Nuzhat, Taha, et al., 2026
Selected projects
Hover a project to run it.
create-rag-app
GitHub ↗ (opens in a new tab)A Python CLI that scaffolds production-ready, Dockerized RAG apps from 13 Jinja2 templates, with 3 retrieval strategies and 8+ integrations.
Bare-metal Prospector
GitHub ↗ (opens in a new tab)A framework-free SDR agent on a raw Python ReAct loop, with a syntax-directed parser that heals tool hallucinations and a three-tier memory.
GPT-2 from scratch
GitHub ↗ (opens in a new tab)Pretrained GPT-2 from scratch in PyTorch, sped up inference with KV caching and speculative decoding, and instruction-tuned it on Alpaca.
Agentic RAG
GitHub ↗ (opens in a new tab)A LangGraph planner and re-planner agent for retrieval, evaluated with RAGAS at over 95% faithfulness.
