AI / RAG Application Developer · Production chatbots and AI agents that ship · Multilingual support, grounded retrieval, no hallucinations · Ex-Samsung R&D · 15+ years of software engineering
AI agents, MCP servers, RAG chatbots and voice agents that answer from the source, not a guess. Grounded in your real documents, with validation layers that refuse to write down what nobody said. Built for real users, real traffic, real outcomes.
I run Vector Forge · Ex-Samsung R&D · 17+ years of software engineering · Based in Dhaka, Bangladesh · Available worldwide remote
Available for direct engagements and for white-label contract work behind agencies and product studios.
MCP servers and agent tools that plug your product, documents or data into AI assistants. The tools are deterministic, the answers cite their source, and where it matters there is no model in the live request path, so responses are fast, predictable and testable. Hosted on AWS (Amazon Bedrock AgentCore, or your own containers), with observability from day one.
Production RAG chatbots over your docs, PDFs, Notion, or SQL — with citations, not hallucinations. That includes the unglamorous ingestion work: messy PDFs, broken text encodings, and schema-checked extraction audited against the source. Multilingual AI agents that handle Bangla, Banglish, English, and other low-resource or script-mixed languages, which makes them especially useful for South Asian, Middle East, and emerging-market audiences.
Voice agents that take calls and fill in structured records — intake, verification, booking, triage — with a validation layer so a mis-heard account number or name never quietly lands in your database. If your agent is going to write to a system of record, that layer is not optional.
Beyond that, I build Messenger, WhatsApp, Telegram, and Slack bots wired to real business data, custom AI copilots embedded inside SaaS products, and the evaluation pipelines, observability, and guardrails that keep all of it from silently regressing in production.
I also handle the cloud infrastructure to keep it running reliably — AWS, Kubernetes, Terraform, CI/CD, CloudWatch, Prometheus, Grafana. One contractor, one accountable line, no hand-off between the AI person and the DevOps person.
A self-hosted MCP server built for Alexa+ that diagnoses home appliance problems from the real manufacturer manual for the appliance you own, and cites the page. Ask about an error code, or describe a symptom in plain words (“my washer has too many suds”), and FixIt returns the manual’s causes and repair steps. If the manual doesn’t say, FixIt says that too. It never invents a repair step or a safety claim.
Built for the Build, Ship, Shape: Amazon Developer Hackathon (Alexa+ track, plus the AWS Builder and Open Source mini challenges), and maintained as an open-source project.
🎬 Watch the demo · 📦 View source (MIT) · 🏁 Devpost
The finding that shaped the whole build: in real conversations, the assistant described an appliance code as “safe” because the manual’s safety-warning list for it was empty. An empty field is not a fact. Now an empty field produces “the manual doesn’t list one”, an unknown code produces “not found”, and both rules are pinned by tests so they can’t quietly regress.
Key decisions:
Stack: Python 3.12 · MCP Python SDK (spec 2025-11-25, Streamable HTTP, MCP Apps) · Amazon Bedrock AgentCore (Runtime, Memory) · Amazon Bedrock (Claude, Nova Pro) · Amazon Polly · CloudWatch · SNS · ECR · Docker · FastAPI · PyMuPDF · pydantic · GitHub Actions
At a glance: 7 real manufacturer manuals · 46 error-code records · 114 symptom rows · 6 MCP tools · 60+ entry friction log for the Amazon developer teams
Built to be contributed to: MIT license, CI, tagged releases, a quickstart that runs without an AWS account, “good first issue” tickets, and a guide for manufacturers who want their manuals supported.
A voice agent that takes insurance claims by phone and cannot write a value into the record unless server-side code approves it. The agent listens and proposes. A validator decides, returning one of three verdicts — accepted, unconfirmed, or rejected — along with the exact sentence the agent then reads back, spelled phonetically.
🎧 Call it yourself · 📊 See what validation catches
The finding that shaped the whole build: I fed the speech recogniser the list of valid policy numbers to improve accuracy. It improved — and it started rewriting mis-heard numbers into real ones. A transcript that began as C411 came back as a genuine policy number belonging to a different customer, and it passed validation perfectly, because the match had been manufactured before the check ever ran.
Key decisions:
BX7-4402 and BX7-4420. Several pin design decisions rather than behaviour, so a later change that undoes one fails and explains why it existed.Stack: Python 3.14 · AssemblyAI Voice Agent API (Universal-3.5 Pro) · FastAPI · raw WebSocket relay · AudioWorklet (PCM16 @ 24 kHz) · Server-sent events · Docker · NGINX · Let’s Encrypt
Recognised by the platform team: three documentation corrections from this build were published by AssemblyAI, and the recogniser-bias finding was escalated to their research team.
A production multilingual RAG chatbot deployed on Facebook Messenger for Minimal Limited, an interior design company in Dhaka. Customers send questions in Bangla, Banglish, or English — the bot always replies in formal Bangla, grounded in a curated knowledge base, with graceful human takeover when confidence is low.
The one decision that paid off most: embed the question, not the answer. Customers send questions, so questions belong in the searchable space. Fixed more “wrong answer” bugs than any prompt tweak.
Key decisions:
Stack: Python 3.13 · OpenAI (text-embedding-3-small, gpt-4o-mini) · FAISS (IndexFlatIP, L2-normalized) · FastAPI · Uvicorn · Facebook Graph API · Pytest
At a glance: 224 curated Q&A entries · 14 intents · top-k=3 retrieval · embedding-dim 1536
Read the full case study · View on GitHub
Paste any error — a Python traceback, a Kubernetes pod crash, a Docker build failure — and get a real fix back. Hit the exact same problem again later, even on a different machine, and it recognizes it and tells you what fixed it last time. Live on the Anna App Store.
The design call that shaped the whole thing: two logs of the same underlying error almost never look byte-identical — timestamps, pod names, and file paths all differ. Strip everything volatile, classify what remains, hash it, and the same problem gets recognized as the same problem no matter how differently it’s phrased each time.
Key decisions:
Stack: Python (stdlib only) · PyInstaller · JSON-RPC · GitHub Actions
Also produced: a reusable app template extracted from this build, so the next developer on the same platform doesn’t have to rediscover the same setup issues.
I’ve shipped real software for 17+ years — not just AI demos.
I bring engineering rigor: evals, logging, retrieval tuning, and guardrails. The unglamorous work that decides whether your AI survives contact with real users.
I keep the model out of the places it doesn’t belong. In FixIt, every live answer is a deterministic lookup into data extracted and audited offline; in the claim intake agent, the layer that decides what gets written contains no AI at all. Models are excellent at proposing. Code should decide.
I also measure before I ship, and I publish the weak numbers alongside the strong ones. Three separate attempts to tune my way out of a speech recognition problem were tested and thrown away because the numbers said they didn’t work. FixIt’s symptom matcher reports its held-out recall of 4 in 10 right next to its tuned results, because a number you only show when it’s flattering isn’t a measurement.
And I don’t trust a metric until I know what it hides. On a recent project, FP16 quantization looked fine on mean error while 15% of gripper commands silently flipped sign — close became open. The average was healthy and the system was broken. Finding that class of failure is most of what reliability work actually is.
And because I can build both the AI and the cloud infrastructure it runs on, there’s no hand-off between the AI person and the DevOps person. One contractor, one accountable line.
If your team sells AI work and needs the engineering layer behind it, I work as a white-label contractor under your brand and under NDA. You keep the client relationship. I stay invisible unless you want me on the call.
Where I usually come in:
Scoped projects or ongoing capacity, whichever suits the engagement.
Agency engagements are usually faster: a short technical call on the work already sold, a scoped estimate, and a start date.
“Sadi is a highly skilled solutions architect, DevOps expert, and technical project manager with a deep understanding of software development, system architecture, and cloud infrastructure. His ability to streamline complex processes, optimize workflows, and enhance system efficiency made him a key asset to Samsung’s R&D initiatives. I highly recommend Md. Shihabuddin Sadi to anyone seeking a dedicated, skilled, and forward-thinking technical leader.”
— Md Elme Focruzaman Razi, Senior Staff Engineer at Samsung R&D Institute Bangladesh (LinkedIn)
pdf-encoding-repair — new open-source library, published on PyPI
Some PDFs extract as text that looks like text but isn’t: a broken font encoding shifts every character by a constant offset, and search, RAG and extraction pipelines downstream quietly ingest the garbage. Nothing errors, so nobody notices.
I hit this on two of the seven manufacturer manuals behind FixIt, each with a different offset. Rather than patch it inside one project, I extracted the detection and repair into a standalone MIT-licensed library, with guards against false positives so clean PDFs are left untouched, CI, and a tagged release. Anyone ingesting PDFs can now reuse the fix instead of rediscovering it.
fixit-mcp — open to contributors, with a friction log for the platform teams
FixIt is MIT licensed and set up for outside contributors: “good first issue” tickets, issue and PR templates, a security policy, an architecture guide, and a guide for appliance manufacturers who want their manuals supported. A quickstart runs the whole thing without an AWS account.
Alongside it is a friction log of 60+ real entries for the Amazon developer teams, each with what was attempted, what happened, a severity, the workaround, and a suggestion. It covers IAM permissions that CreateAgentRuntime needs but the docs don’t list, a container guide that doesn’t fit MCP servers, an undocumented Alexa+ session model, and the discovery that Claude models on Bedrock are billed through AWS Marketplace, outside promotional credits.
anna-developer-docs — Documentation corrections, merged
While building Error Journal on Anna’s platform, I lost a day to behaviour that contradicted the documentation. Rather than work around it, I traced each discrepancy through the runtime source and wrote up seven findings with replacement text.
All seven were verified as accurate. Six were merged into the public developer documentation, including a capability string that no longer existed in the runtime, a required manifest field missing from the reference table, and a config schema documented with the wrong data type.
The seventh turned out to be a production bug rather than a docs error — a storage-token issue that took the platform team a proper investigation to root cause, traced to a resource silently resetting its visibility when edited through the web UI. Confirmed and fixed in the following release.
“One of the best community write-ups we’ve received — seven precise findings, each verified against actual runtime behavior. We verified all seven items and every single one was accurate.” — platform engineering team
Why it’s here: most of these were found by reading the runtime source rather than re-reading the docs. That habit — verifying behaviour instead of trusting documentation — is the same one that keeps production systems debuggable.
claim-intake-agent — findings published by AssemblyAI
The same habit, on a different platform. Building the voice agent surfaced three places where the published message-sequence and browser-integration docs disagreed with the machine-readable API schema — a field name for reply audio, a field name for agent transcripts, and the identifier on tool results. Each was verified against the live API rather than assumed, reported, and published as documentation corrections.
A fourth finding was a model behaviour rather than a docs error: biasing transcription toward a list of known values can rewrite a mis-heard value into one of them, which is dangerous anywhere the transcript is the evidence being validated. That one was escalated to their research team.
anna-app-template — reusable scaffold, published for other builders
After shipping Error Journal on the same platform, I extracted the parts worth reusing so the next person doesn’t repeat the discovery. JSON-RPC transport with a forward queue for concurrent reverse-RPC calls, persistent storage and model sampling that degrade gracefully when unavailable, three-platform binary CI, and a publish runbook documenting the failure mode at each step.
Clone, run the rename script, get a running plugin — verified from a clean clone rather than assumed.
A selection of supporting projects across AI research engineering, cloud infrastructure, DevOps automation, and software engineering.
Counterexample — PR Verification on IBM Bob 2.0 — evidence over opinions, not tests that pass A pull-request verification tool for AI-generated code, where green CI is a weak signal because the same reasoning that wrote a bug often wrote the tests around it. It runs two independent checks and merges them into one Review Receipt: diff-scoped mutation testing (an AST mutator that injects bugs only into the lines a PR changed, then runs the PR’s own tests against each mutant), and claim falsification, where IBM Bob extracts the concrete claims a PR makes from its diff and linked issue, spawns one subagent per claim, and has each write and run an adversarial test to break it. A GitHub Action posts the mutation layer’s result on every pull request.
Three findings worth the click. On a real test PR, mutation testing scored a reproducible 100% while the PR was still wrong — the bug computes a discount from the wrong variable, which none of the four mutation operator families can express. Claim falsification caught it: 4 of 5 claims falsified, each with a failing test and real output, including a silent breaking change for existing callers that was never planted. The tool had a bug of its own — in Python’s pathlib, joining a temp directory with an absolute path silently discards the temp directory, so concurrent mutant writes landed on the real repo instead of an isolated copy. It surfaced as unexplained working-tree corruption and is now guarded by a loud error. And CI reported 0% for the wrong reason: pytest was collecting unrelated files from GitHub’s merge-commit checkout, so every mutant errored identically. The receipt now shows per-mutant error detail, and the PR comment states plainly that it covers mutation testing, not requirement-level verification.
On the Bob side: a custom mode with scoped tool permissions, two custom skills, and five parallel subagents in isolated contexts. The first full review cost 1.71 Bobcoins. See it on a real pull request.
Tech: Python 3.11 · IBM Bob 2.0 (custom modes, skills, subagents) · AST mutation testing · pytest · GitHub Actions
Bimanual VLA — Table Setting in Simulation — measurement over assumption Two simulated SO-101 arms set a table in MuJoCo: a scripted expert picks four props from a randomized layout, hands a prop between arms when no single arm can reach both the prop and its slot, records the successes as a LeRobot v3.0 dataset, trains an ACT policy on it, and converts the checkpoint to OpenVINO IR for Intel inference hardware.
The engineering interest isn’t the robotics — it’s the discipline. Every design decision traces to a measurement, and the README opens with a table of what was measured and what explicitly was not.
Three findings worth the click. FP16 quantization looked healthy on mean error while 14.94% of gripper commands flipped sign — close became open, which drops whatever the arm is holding; INT8 flips 0.92% and is 3.46× smaller. The original parity check passed and was wrong, because it ran against synthetic noise where a badly wrong precision looks exact; on real frames the same model was off by three orders of magnitude more. And the underperforming policy was handled as a controlled experiment rather than a result to bury: three causes diagnosed, one isolated by removing an image-task confound from the training data, attention measurably redirected (non-plate target contact 0/30 → 7/30) while task competence stayed flat, exactly as the two untouched causes predict.
Tech: Python 3.11 · MuJoCo · LeRobot 0.4.4 (ACT, 51.6M params) · PyTorch · OpenVINO + NNCF · MiniLM-L6
Single-Node Kubernetes Cluster Multi-service web app (React, Node.js, MongoDB) deployed on a single-node Kubernetes cluster using Minikube — Deployments, Services, Ingress, ConfigMaps, Secrets, PV/PVC. Tech: Kubernetes · Docker · Minikube · NGINX Ingress
Kubernetes on AWS (EKS) End-to-end CI/CD on AWS EKS — Fargate, eksctl, Jenkins, DockerHub, and ECR integrations. Tech: AWS · EKS · Fargate · Jenkins · Docker
Terraform IaC Infrastructure as Code patterns for repeatable, auditable cloud deployments. Tech: Terraform · AWS
Prometheus + Grafana Monitoring Monitoring and observability setup for cloud-native applications. Tech: Prometheus · Grafana
Ansible Automation Configuration management and infrastructure automation playbooks. Tech: Ansible · Playbooks
See all repositories on GitHub
17+ years of software engineering across embedded systems, mobile, full-stack, cloud, and AI.
I started at Samsung R&D Bangladesh, where I worked on firmware for handsets shipped across Middle East, Africa, and Bangladesh — including the Bengali Calendar for the Bangladesh region and language support for Swahili, Yoruba, Igbo, Hausa, and Amharic on Samsung feature phones used by millions. That’s where multilingual production software became muscle memory, which weirdly turned out to be great prep for the multilingual RAG work I do now.
After Samsung, I co-founded Training Pool, Bangladesh’s first online training marketplace and SaaS platform. Took it from idea to live product with paying users. Before that, I ran a small dev studio building Android multiplayer games and Bangladesh client projects.
These days I run Vector Forge, shipping production RAG applications, AI agents, MCP servers, and voice agents for founders, agencies, and mid-market teams.
See full work history on LinkedIn
Read my latest posts · RSS feed
If your chatbot is hallucinating, your voice agent is writing down things nobody said, your AI feature isn’t making it past the demo stage, or you want to add a real RAG system to your product without it embarrassing you in front of customers — let’s talk.
Want your product or your documentation inside AI assistants, through an MCP server that answers from the source and says so when it doesn’t know? Same conversation.
Running an agency with AI work sold and no one to build the reliable version of it? That conversation is even shorter.