Syed Taha, AI engineer & researcher

PKT · Karachi

I’m obsessed with building the best agents, extracting every ounce of accuracy under realistic latency and token budgets, backed by extensive trace evals.

About

Turning research on efficient agents into products people can rely on.

Experience

Hover a row for the detail.

Engineering

Résumé ↓
  • 01Brainbox AutomationsAI EngineerDec 2025 – NowDeployed a safety-guardrailed AI coaching backend on GCP Cloud Run (FastAPI, Vertex AI) serving 300+ minor and adult athletes
  • 02Epistemy UKSoftware Engineer (AI) & Team LeadSep – Dec 2025Led the end-to-end build of an AI tutoring platform with a Nest.js backend (20+ endpoints) and a Next.js frontend
  • 03CogniMind AIMachine Learning InternFeb – Apr 2025Raised VLM extraction accuracy by 20% through prompt engineering
  • 04RapidsAI Machine Learning InternSep – Dec 2024Built a RAG chatbot (Streamlit, FastAPI) with contextual session management

Research

Academic CV ↓
  • 01IPT Lab, NUSTResearcherJun 2026 – NowLeading 2 studies on redundancy and efficient inference in foundation models
  • 02Bradbury LabResearch Intern · remoteApr 2025 – Jan 2026Proposed a training-free layer-merging method based on Tucker decomposition to cut parameter count
  • 03MachVis LabUndergraduate Research InternOct – Dec 2025Built an LLM-based factual-verification framework for the clinical accuracy of generated pathology reports, beyond BLEU/ROUGE
  • 04NUST · remoteUndergraduate Research InternJun – Sep 2025Combined MedSAM with meta-learning for few-shot dental radiograph segmentation, +12% on scarce disease classes

Inside the research

Measuring structural redundancy against proper controls, compressing models, and attaching statistical guarantees to compression decisions. Open to exploring agentic research on resource-constrained devices.

Similarity ≠ importance

Blocks of similar layers look removable. Similarity predicts which ones no better than an untrained network.

Certified early exit

Each band is an input through the layers, stopping where it is confident enough, with error ≤ α at 95% confidence.

Distillation · LiteDoc

A large teacher distilled through a narrow channel into a small task-specific student.

Facts, not words · MedGemma 1.5 4B

MedGemma 1.5 4B reads a slide into a fluent report, but gets treatment-critical facts wrong 41–67% of the time, and ROUGE-L can’t tell.

Encoder compute saved by certified early exit
14–39%
Spearman’s rho between similarity-based removal order and an untrained network’s
≥0.88
Of the DeepSeek-VL2 (MoE) teacher’s performance retained by LiteDoc, on average
89.6%
  1. C1

    LiteDoc: Distilling Large Document Models into Efficient Task-Specific Encoders

    Tayyab, Taha, Adrian, Ulrich, Momina, Faisal, ICDAR 2026, Springer LNCS

    DOI ↗ (opens in a new tab)
  2. M1

    Similarity Is Not Importance: On Measuring Representational Redundancy in Wireless Foundation Models

    Taha, Draft on request

    In prep.
  3. M2

    Certified Early Exit in a Wireless Foundation Model When the Signal-to-Noise Ratio Must Be Estimated

    Taha, Draft on request

    In prep.
  4. M3

    Accuracy of Craniometric Features in Gender Estimation Using Machine Learning Algorithms on University of Tennessee (UT) and Howells Datasets

    Nuzhat, Taha, et al., 2026

    Submitted

Selected projects

Hover a project to run it.