Skip to content
Sailesh Dahal

Start a conversation

Tell me what you are trying to do. If it is a bad idea I will say so, and if there is a cheaper way to get the same result without a model in the loop, I will say that too. No pitch deck, no discovery call before there is anything to discover.

Goes straight to my inbox, with no third party in between and no spam screen to solve. I read every one myself.

Rather not use a form? Email hi@saileshdahal.com.np or find me on LinkedIn and X.

What I can help you build

Example briefs, ordered from the cheapest useful thing to the deep end. Not a list of past projects. The work I have actually shipped is on the home page, with links you can open.

Start here

  • AI where it already fits

    Summarise, classify, extract, draft. Wire a model into the workflow your team already runs (a support macro, a triage queue, a report that writes its own first draft) with a prompt you can read and a test that tells you when it drifts.

  • Ask your documents

    Retrieval over your own corpus: contracts, policies, tickets, wiki, PDFs. Chunking and embeddings tuned to your documents, hybrid search with a reranker, and every answer citing the passage it came from so a reader can disagree with it.

Production

  • RAG that survives contact with real data

    The version that holds up past the demo: incremental ingestion, freshness and permissions carried through the index, query rewriting, retrieval evals with real recall numbers, and a fallback that says "I do not know" instead of inventing a clause.

  • Agents with hands

    Agents that act on your systems through typed tools and MCP servers (book, refund, update, escalate) with scoped permissions, idempotent actions, a full audit trail, and a kill switch. Human approval on anything expensive or irreversible.

  • Evals, guardrails and the release gate

    A graded dataset built from your real traffic, an eval suite in CI that blocks a prompt or model change from shipping when quality drops, plus injection and PII checks, refusal behaviour, and per-request cost and latency budgets you can watch.

Deep end

  • Your own inference stack

    When the API bill or the data boundary stops making sense: open-weight models served on your own GPUs, quantisation and batching sized to your traffic, LoRA fine-tunes on your data, and a routing layer that sends the easy 80% to the small model. The fine-tuning and quantisation half of this is thesis work I can show; no specific serving stack is named here because that choice depends on your hardware.

  • Multimodal and voice

    Vision, audio and documents in the same pipeline: read the invoice, watch the video, hear the call. Streaming speech in and out for a phone or in-app agent, with latency budgets that keep a conversation feeling like a conversation.

  • AI inside a mobile app

    The part most teams underestimate: streaming into a React Native or Flutter UI, on-device and hybrid inference, offline behaviour, token cost per session, and a review-safe story for what the model can and cannot do on a user’s phone.

None of these quite it? Describe it in your own words. The form takes free text for a reason.