Zain Raisan

AI engineer in Abu Dhabi, mostly drawn to high performance computing. I take models to the hardware they actually run on, GPUs, NPUs and agents in production.

Open to AI engineering roles · 30 days notice CV Email GitHub LinkedIn
How I think about speed

Most of my work is making sure the GPU never waits.

Decode, preprocess, infer, encode. Run them one after another and three engines sit idle while the fourth works. Overlap them, then batch the expensive one, and the same card does several times the work. Switch the schedule and watch the engines.

throughput0fps
latency0ms
speedup1.0×
timeline · last 60 ms of simulated time illustrative stage costs · decode 3 · pre 2 · infer 6 · encode 2.5 ms

01 / 06Computer vision · open source

Argus

Thirty two camera streams, one laptop GPU.

perf32 streams · 1 GPU · 8 GB · decode to encode on device

A C++20 detection pipeline. FFmpeg decode, TensorRT YOLO, overlay and NVENC encode, each in its own queue stage, keeping the video on the GPU from the first byte to the last.

32concurrent streams of YOLO11s on one 8 GB RTX 3070 Laptop, as demonstrated
5input and output protocols, from RTSP and HLS to MP4 and Matroska

Real recording from the repository

02 / 06Android · Google Play

Orb

A vision model in your pocket, in airplane mode.

perfNPU · GPU · CPU picked at runtime · 0 network calls

An offline assistant running quantized vision language models on the phone. It picks the Snapdragon NPU, GPU or CPU at runtime, transcribes with Whisper on device, and sends nothing anywhere. The inference layer underneath is open source as native-onnx.

4quantized VLMs, Qwen3-VL 2B and 4B, MiniCPM-V 4.6 and Gemma 4 E2B
5execution backends in native-onnx, from CPU to the Hexagon NPU
Orb screen 1
Orb screen 2
Orb screen 3

03 / 06Agents · open source

Deputy

Any web page, as a typed API for agents.

perf216,888 → 133,366 tokens · 26.1 s → 17.2 s wall clock

A Chrome extension and MCP server that grafts the W3C WebMCP standard onto sites that never adopted it. Four attributes on the page's own form, and Chromium generates the schema and runs the submission. Forms that change state are filled and left for a human to submit.

1.63×fewer tokens than Claude driving Playwright MCP on the same form, in my benchmark
115tokens per page observation, against 14,170 for an accessibility snapshot

Real recording from the repository

04 / 06Python · PyPI

globalmm

Teaching a local model to see, without training it.

perf1 closed form solve · 1152 × d · seconds, not GPU hours

Generates a llama.cpp compatible vision projector for any local LLM. Instead of hours of gradient descent on caption data, it solves for one matrix in closed form, mapping frozen SigLIP patches into the model's own embedding space.

1least squares solve, a 1152 by d matrix, in seconds
81soft tokens per image in the target model's embedding space

W = (X⊤X)−1 X⊤ Y

SigLIP patches · 1152 W one solve 81 soft tokens · d

05 / 06Product · live

Wirecopy

A daily briefing you can actually finish.

perf~2,500 feeds · one edition a day · inside the free CI tier

AI releases, security events and research for developers, built and run end to end. One Go binary as server or worker, Postgres with pgvector, and a Cloudflare edge read model, published every morning on GitHub Actions inside the free tier.

2.5k+feeds across 20 source families, licence checked per source in the pipeline
0cookies used to count reads. Nonce only CSP and SSRF refusal in the fetcher
Wirecopy screenshot

06 / 06Agents · private build

Sourcing Agent

Eight workflows that are not allowed to sign anything.

perf376 tests · 45 evals · 650 MiB of RAM given back

A governed procurement platform I built alone. Agents draft, compare and route, while approvals run as durable human in the loop workflows on Temporal. They cannot approve spend, reject a supplier or sign a contract on their own, and every call goes through a gateway the customer controls.

376backend tests, plus a 45 case offline evaluation suite in Arabic and English
650 MiBof container memory saved by sharing model weights across workers
request 8 agentworkflows retrievalQdrant + BM25 human gateTemporal audit logevery event action gateway: every LLM call through a customer controlled endpoint

Mostly, I chase
performance.

High performance computing is the part of AI I like most. Fitting models onto the hardware that exists, keeping data on the device, and measuring before claiming. These are the numbers from the projects on this page.

zr-smi · performance logmeasured or demonstrated, never extrapolated
procmetricvaluecontext
argus concurrent streams, YOLO11s 32 one RTX 3070 Laptop, 8 GB, as demonstrated
argus video path GPU CUVID decode, TensorRT, NVENC encode
haykal real time factor, Spark-TTS 0.5B 0.95 RTX 3070 Laptop, 8 GB
haykal real time factor, F5-TTS on ONNX 0.43 prototype, not shipped
haykal first mouth movement 366 ms down from 516 ms, best case
sourcing container memory returned 650 MiB shared weights via copy on write
legal-rag response time < 2 s Qwen embedding and reranker on Triton
deputy tokens per page observation 115 against 14,170 for an accessibility tree
globalmm projector fit 1 solve closed form, no gradient descent

In production
for Abu Dhabi government

At the Department of Municipality and Transport since 2025. The code is not public, so these describe what the systems do, at the level I can share.

< 2 s

Bilingual legal retrieval

600+ Arabic and English legal documents, one vector store per language, self hosted Qwen embedding and reranker models on NVIDIA Triton.

13

Agent analytics platform

Thirteen specialised agents over two government databases, serving visual analytics to 10,000+ internal users.

1,000+

City wide road defects

A quantized Qwen3 VLM with an LLM verification layer, reading dashcam and CCTV footage across the city.

1,000+

Live video analytics

YOLOv11l with TensorRT on CCTV and RTSP sources, republishing annotated streams.

Lead

Demand onboarding platformlead developer

A conversational planner that elicits requirements, finds reusable work in existing contracts, and estimates budget and timeline before anything is built.

MCP

API catalogue standardinitiated and led

One OpenAPI standard across teams behind an automated lint gate, with conforming APIs exposable to agents as MCP servers.

20+

Model serving platform

LLMs on vLLM, embeddings on Triton, 20+ containerised AI apps on AKS, observed with Grafana, Prometheus, Loki and Tempo.

300+

AI Champions platformprimary developer

Courses, admin workflows and chat over governed model endpoints for 300+ government employees.

Experience

2025 to now

AI Engineer

Department of Municipality and Transport, Abu Dhabi

Production RAG, agents and computer vision. Lead developer of the demand onboarding platform.

2024

Technology Consultant Intern

Deloitte, Dubai

Clustered regulations with a T5 model and built a Three.js network topology viewer.

2024

BS Computer Science and Computer Engineering

American University of Sharjah

GPA 3.81. Published in IEEE Access.

2023

Software Engineering Intern

Mohammed Bin Rashid University of Medicine, Dubai

Automated a document pipeline, cutting manual processing time by 80%.

2023

Embedded Systems Intern

Remal IoT, Sharjah

Designed a PCB shield and ran an ESP32 workshop for 30+ students.

Publication

Dental Radiograph Analysis for Improved Diagnosis · IEEE Access, 2024

Contact

Let's build
the hard part.

AI engineering, forward deployed and ML infrastructure roles. Abu Dhabi based on a UAE Golden Visa, open to remote, 30 days notice.