Abu Dhabi, UAEPh.D. Researcher, AI & Machine Learning

Malcom Mudhungwaza

I build LLM systems that hold up in production, and research where they fail.

Agentic AI, RAG, evals and LLM serving — designed end to end and measured against what the system is actually for. Alongside it, doctoral research on guardrails: where safety alignment breaks down across language and modality, and whether the benchmarks measuring it test refusal or reading ability.

  • Senior AI Engineer
    GenAI, LLM & agentic systems
  • Ph.D. Researcher
    AI safety & alignment
  • 5+ years
    production LLM & ML systems
Malcom Mudhungwaza

Malcom Mudhungwaza

Senior AI Engineer · Abu Dhabi, UAE

Core capability

AI agents, RAG, software engineering and platforms

All five domains →

What I build

Systems, not demos

The things I am usually brought in to design and ship. Each has a version running against real traffic, and a version published here as open research.

  • Agentic document pipelines

    Multimodal extraction into structured output, LLM reasoning over the result, confidence gating per document class, and a human review path for the cases that should not be automated.

  • RAG systems

    Hybrid dense and lexical search with re-ranking, chunking chosen against the document shape, and retrieval quality measured separately from generation quality.

  • LLM serving & inference optimisation

    Quantised models behind an inference gateway, sized against a stated latency target and a cost per million tokens, with routing and per-tenant budgets.

  • Post-training & domain adaptation

    Supervised fine-tuning and preference optimisation on open-weights models, with the objective chosen against the feedback actually available rather than the newest paper.

  • Evals & CI regression gates

    Golden datasets, LLM-as-judge scoring calibrated against human labels, and regression gates that fail a build when quality drops by a detectable margin.

  • Voice agents & multimodal

    Streaming ASR into dialogue orchestration into speech synthesis, budgeted per hop against time-to-first-audio rather than total generation time.

  • Agentic AI & MCP servers

    MCP servers exposing internal systems as tools, orchestration with checkpointing that survives a restart, and topology chosen from the failure mode that matters most.

  • Platform layers & observability

    Async services with bounded concurrency, retry classification at the transport boundary, caching with a measured hit rate, and schemas read against their query plans.

Research

A doctorate on where safety alignment breaks

Ph.D., Artificial Intelligence & Machine Learning, in progress. The dissertation asks a question that current benchmarks cannot answer:

How much of the apparent cross-modal safety gap is a reading-ability artifact rather than an alignment gap?

A model that cannot read a request scores the same as one that read it and refused. Separating the two is the contribution.

Field
AI safety and alignment
Working papers
Multi-Agent Orchestration with Model Context Protocol: A Framework for Reliable Tool-Augmented LLM SystemsPreprint in preparation · arXiv, 2026Safety Alignment Across Language and Modality: Perception Confounds in Multimodal LLM Guardrail TransferPreprint in preparation · arXiv, 2026
Open research
26 repositories · 950 tests

Built with

  • PyTorch
  • Transformers
  • TRL
  • vLLM
  • AutoAWQ
  • LangGraph
  • FastAPI
  • PostgreSQL
  • pgvector
  • Redis
  • FAISS
  • sentence-transformers
  • rank_bm25
  • ONNX Runtime
  • faster-whisper
  • piper
  • OpenCV
  • Ultralytics
  • NumPy
  • SciPy
  • Matplotlib
  • Pydantic
  • httpx
  • anyio
  • arq
  • OpenTelemetry
  • pytest
  • datasketch
  • fastText
  • C++
  • Python
  • WebSockets

Contact

Get in touch

Happy to talk about anything here, or about LLM systems, evaluation and serving generally. Mentioning a repository by name gets you a faster and more useful answer.

Based in
Abu Dhabi, UAE