Blog

Writing on agentic AI, blockchain, and AI systems architecture.

Technical writing by Mudassir Khan — agentic AI consultant and AI systems architect. New posts published directly on this site. Earlier writing on Dev.to.

RAG & AI Engineering

Production RAG pipelines, vector database selection, fine-tuning tradeoffs, and LLM governance.

Cover illustration for: JavaScript Programming: Complete Beginner Guide
AI Engineering15 min read

JavaScript Programming: Complete Beginner Guide

JavaScript is the language of the web and modern AI tooling. This guide covers variables, functions, loops, objects, async, and your first program.

Read post →
Cover illustration for: Production RAG: Why Retrieval Fails and How to Fix It
RAGAI Engineering11 min read

Production RAG: Why Retrieval Fails and How to Fix It

Most RAG failures happen at retrieval, not at the LLM. Chunking strategies, hybrid search, reranking, and RAGAS metrics for production RAG pipelines

Read post →
Cover illustration for: RAG vs Fine-Tuning: The Decision Guide for Production
RAGLLMs9 min read

RAG vs Fine-Tuning: The Decision Guide for Production

RAG vs fine-tuning — which approach fits your production LLM? Covers RAG vs fine-tuning vs prompt engineering, LLM fine-tuning vs RAG, when to use

Read post →
Cover illustration for: pgvector vs Pinecone vs Weaviate: How to Choose
RAGAI Engineering9 min read

pgvector vs Pinecone vs Weaviate: How to Choose

Start on pgvector, migrate when you must. An AI architect's guide to choosing a vector database — when each option wins, what the performance numbers

Read post →
Cover illustration for: AI Governance Framework for Production LLMs: The Checklist
AI SystemsAI Engineering9 min read

AI Governance Framework for Production LLMs: The Checklist

A practical AI governance framework for production LLM systems — the five layers every team needs, the EU AI Act obligations coming into force

Read post →

More posts coming soon

A 30-post technical series on agentic AI architecture, LangGraph patterns, blockchain compliance engineering, and production AI systems is in progress. Follow on Dev.to or LinkedIn to be notified.

FAQ

Frequently asked questions

When should I hire an agentic AI consultant?

Hire an agentic AI consultant when your team needs to build autonomous, multi-step AI workflows — such as agents that browse the web, call APIs, or orchestrate multi-model pipelines — and lacks in-house experience with LangGraph, tool-use reliability, or production agent safety. A consultant is faster than hiring a full-time engineer if your timeline is under 6 months.

What is the difference between RAG and fine-tuning?

RAG (Retrieval-Augmented Generation) fixes knowledge gaps by injecting relevant documents at inference time — ideal for private or frequently updated data. Fine-tuning fixes behaviour gaps by adjusting model weights on curated examples — best when you need consistent tone, output format, or domain-specific reasoning. In 2026, most production systems use both: a fine-tuned model as the base, with RAG for dynamic knowledge retrieval.

What does an AI systems architect do?

An AI systems architect designs the end-to-end infrastructure for AI products — including model selection, data pipelines, vector stores, agent orchestration, evaluation frameworks, and deployment strategy. Unlike an ML engineer who focuses on model training, or a data scientist focused on analysis, an AI systems architect owns the full production system design.

What is the difference between Ethereum and Solana?

Ethereum uses the Gasper consensus mechanism (Proof of Stake + Casper FFG), prioritising security and decentralisation over speed — it processes roughly 15–30 TPS. Solana uses Proof of History combined with Tower BFT, achieving 2,000–65,000 TPS at lower cost, but with greater centralisation risk. Ethereum is preferred for DeFi and high-value applications; Solana suits high-frequency, low-cost use cases.

What are multi-agent design patterns in production AI?

The four multi-agent patterns that work reliably in production are: Sequential (agents chain output to input in a fixed order), Orchestrator-Worker (a coordinator agent delegates tasks to specialised sub-agents), Hierarchical (nested orchestration across multiple management levels), and Dynamic Handoff (agents route tasks to the best-suited agent at runtime). LangGraph is the most common framework for implementing all four.

Should I hire an AI engineer full-time or use a consultant?

Use a consultant if you need a working system in 1–3 months, have a defined scope, or are pre-Series A. Hire full-time when you have a live product requiring ongoing iteration, a team to mentor, and 6+ months of runway budgeted for the role. The hybrid model — consultant to build the initial system, then a mid-level engineer to maintain it — is what most seed-stage teams in 2026 actually use.

Blog — Agentic AI, Blockchain & AI Architecture Writing