Puja Set NepalSource on GitHub

Documentation

Everything needed to call Jaynepal 1.1.

Jaynepal 1.1 is a Nepali-first language model made in Nepal and served by Puja Set Nepal from a hosted, OpenAI-compatible endpoint. This page covers what the model is, how to send it a request, and what it is genuinely not good at yet.

Overview

You reach Jaynepal 1.1 over plain HTTP. There is no SDK to install, no weights to download and no GPU to rent: send a POST to the completions endpoint with your messages, and the answer comes back — streamed by default.

The interface is the OpenAI chat-completions shape, so an existing client library or agent framework works by changing two values: the base URL, and the model id.

Base URL     https://pujasetnepal.com/v1
Model id     jaynepal-1.1
Endpoint     POST https://pujasetnepal.com/v1/chat/completions
Models       GET  https://pujasetnepal.com/v1/models

How a request is served

your app  ──►  hosted API  ──►  Jaynepal 1.1  ──►  streamed reply
                (pujasetnepal.com)   (9B)

The API is the boundary you integrate against. Which host the model runs on, how it is batched and how it is scaled are ours to change without asking you to redeploy, as long as the contract below holds.

Model card

FieldValue
ModelJaynepal 1.1
Model idjaynepal-1.1
Version1.1
Parameters9B
LanguagesNepali (Devanagari) · Romanised Nepali · English
Context window8,192 tokens
Serving interfaceOpenAI-compatible HTTP API, SSE streaming
Trained bySamir Puri — Bharatpur, Chitwan, Nepal

What it is built for

Nepali conversation and Nepali prose. It replies in the language you wrote in — Devanagari, Romanised Nepali or English — and holds that choice instead of drifting back to English. It handles the ordinary work a Nepali-speaking user needs from an assistant: explaining concepts, drafting text, translating between the three registers, and writing code with English technical terms around Nepali explanation.

It also writes Markdown, because structure is what makes a long answer readable — headings, lists, tables and fenced code blocks come back as Markdown rather than as plain paragraphs.

What it is not

Stated plainly, because these are the questions a reader will actually hit:

LimitationWhat that means in practice
It is small9B parameters, not a frontier model. It is strong for its size in Nepali and weaker than a very large model on hard reasoning, long maths and obscure world knowledge.
Academic Nepali is unevenFormal legal, medical and academic registers are less reliable than everyday conversation. For anything published, have a native speaker review the output.
It is not a knowledge baseIt can state things confidently and wrongly. Treat factual claims — especially dates, numbers and current events — as drafts to verify.
No tool useIt does not browse, call functions or run code. It answers from the prompt and what it learned while training.
Devanagari is not transliterated for youAsk for Romanised Nepali or Devanagari explicitly if the client or channel needs one of them.

Quickstart

Nothing to install. Export the base URL and your key, then send a request.

export JAYNEPAL_API="https://pujasetnepal.com/v1"
export JAYNEPAL_API_KEY="..."      # when access is keyed

curl -N "$JAYNEPAL_API/chat/completions" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $JAYNEPAL_API_KEY" \
  -d '{
        "model": "jaynepal-1.1",
        "messages": [
          { "role": "user", "content": "लमजुङ किन प्रसिद्ध छ?" }
        ],
        "stream": true
      }'

Confirm the model is there

curl "$JAYNEPAL_API/models"

{ "object": "list",
  "data": [ { "id": "jaynepal-1.1", "object": "model" } ] }

From Python, with an OpenAI client

from openai import OpenAI

client = OpenAI(base_url=os.environ["JAYNEPAL_API_URL"],
                api_key=os.environ["JAYNEPAL_API_KEY"])

stream = client.chat.completions.create(
    model="jaynepal-1.1",
    messages=[{"role": "user", "content": "नमस्ते, तपाईं कस्तो छ?"}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
Any OpenAI-compatible client works — the base URL and the model id are the only two things that change.

API reference

POST /chat/completions

POST https://pujasetnepal.com/v1/chat/completions
Content-Type: application/json
Authorization: Bearer <key>        # when access is keyed

{
  "model": "jaynepal-1.1",
  "messages": [
    { "role": "system",    "content": "optional persona" },
    { "role": "user",      "content": "नमस्ते" }
  ],
  "temperature": 0.7,
  "max_tokens": 1024,
  "stream": true
}

Request body

FieldTypeDefaultNotes
modelstringjaynepal-1.1Required. The only id currently served.
messagesarray—Required. Roles are user, assistant and system. Send the whole conversation each time; the model holds no state between requests.
temperaturenumber0.70 is nearly deterministic, 1 is loose. Lower it for factual or technical answers, raise it for drafting.
max_tokensinteger1024Ceiling on the reply. Generous values cost latency you do not get back if the model finishes early.
streambooleantrueServer-sent events instead of one JSON body. See Streaming below.

Response — non-streaming

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "model": "jaynepal-1.1",
  "choices": [
    {
      "index": 0,
      "finish_reason": "stop",
      "message": { "role": "assistant", "content": "..." }
    }
  ],
  "usage": { "prompt_tokens": 18, "completion_tokens": 96, "total_tokens": 114 }
}

GET /models

Lists what the endpoint is serving. Useful as a cheap liveness check, and as the first thing to try when a request fails.

GET https://pujasetnepal.com/v1/models

{ "object": "list", "data": [ { "id": "jaynepal-1.1", "object": "model" } ] }

Streaming

With stream: true — the default — the response is a sequence of server-sent events, each carrying one fragment of the reply, terminated by a literal [DONE] line.

data: {"choices":[{"delta":{"content":"लमजुङ"}}]}
data: {"choices":[{"delta":{"content":" जिल्ला"}}]}
data: {"choices":[{"delta":{"content":" हो"}}]}
data: [DONE]

Consuming it without breaking Nepali

Read the response as a byte stream and decode with buffering enabled, splitting lines only on complete newlines. A Devanagari character is several bytes, and a token can land in the middle of one; decoding each chunk independently turns that into mojibake in exactly the language this model exists to serve.

reader = response.iter_lines(decode_unicode=False)
buffer = b""

for chunk in reader:
    buffer += chunk + b"\n"
    lines = buffer.split(b"\n")
    buffer = lines.pop()          # keep the partial line for next time
    for line in lines:
        if not line.startswith(b"data: "):
            continue
        payload = line[6:]
        if payload == b"[DONE]":
            break
        text = json.loads(payload)["choices"][0]["delta"].get("content", "")
        print(text, end="", flush=True)
Fragment boundaries are not word boundaries. Accumulate deltas and re-render; never assume one chunk is one word or one character.

Errors

Failures are JSON, not an empty stream — so a client can tell “the model said nothing” apart from “the model could not be reached”.

{
  "error": {
    "message": "human-readable description",
    "type": "invalid_request_error",
    "code": "bad_request"
  }
}
StatusCodeCause
400bad_requestBody was not JSON, messages was empty, or a message had no valid role and string content.
401invalid_api_keyMissing or wrong key, when the endpoint is keyed.
404model_not_foundThe model field does not name the served model.
429rate_limitedToo many requests in flight. Back off and retry.
502upstream_unreachableThe model host could not be reached — usually the serving session not running.
503unavailableThe model is loading. Retry shortly.
504timeoutNo response within the gateway window. Long answers on a CPU-only host can hit this.
Retry 429 and 503 with exponential backoff. Do not retry 400 — the request itself is wrong and will fail again.

Access and status

The canonical production base URL is https://pujasetnepal.com/v1. Access is arranged directly at the moment rather than granted by a signup form, and the endpoint you are given is the one to use in base_url.

The hosted endpoint is not yet publicly attached. Today the model runs behind a session-bound GPU host, which means its address changes when the session restarts. That is fine for evaluation and integration work, and not yet fine for a product that promises five nines. A permanently hosted endpoint is the first item on the roadmap below.

Bringing up an endpoint yourself

The interface is standard, so any server that speaks OpenAI chat completions can front the model while the hosted one is being brought up — a rented GPU box with vLLM, llama.cpp, or a GPU Space. Point your client at it and the code below does not change.

# 1. serve it
vllm serve <jaynepal-1.1> --port 8000 --max-model-len 8192

# 2. confirm the contract
curl http://localhost:8000/v1/models

# 3. use it — the only change your client needs
export JAYNEPAL_API=http://localhost:8000/v1

Limits and roadmap

The gaps, stated as gaps. A reader deciding whether to build on a model needs these more than it needs a feature list.

TodayConsequence
No self-serve keysAccess is arranged directly, so onboarding is not instant.
The evaluation endpoint is session-boundIt is offline between GPU sessions, and its address moves.
No published benchmarkQuality claims are qualitative until the evaluation below ships.
8,192 tokens contextVery long documents must be chunked or summarised rather than pasted whole.
No tool callingIt cannot browse, run code or call your functions.

Roadmap

  • A permanently hosted endpoint on a persistent GPU host.
  • Self-serve API keys with per-key quotas and usage reporting.
  • A published Nepali evaluation — comprehension, generation and Romanised transcription — against comparable models.
  • A quantised release for people who want to run it on their own hardware.
  • Deeper register coverage for formal, legal and academic Nepali.

Credits

Jaynepal 1.1 and this platform are built by Samir Puri (DevSamirX) in Bharatpur, Chitwan, Nepal.

The model is published on Hugging Face. The source for this website is on GitHub under the MIT licence, and issues and pull requests are welcome.