KontyraKontyra Docs
Neural Engine•NanoKora Architecture

KORA (NanoKora Foundation Model)

The complete technical specification for Kora: modern decoder-only transformer architecture, RMSNorm, Rotary Position Embeddings (RoPE), SwiGLU feed-forward networks, and low-rank LoRA adaptation.

1. Overview & NanoKora Engine

KORA is Kontyra's native intelligence assistant, constructed from first principles to demystify modern large language model engineering while delivering fast, deterministic code assistance embedded directly within DevOS and Kontyra APIs.

NanoKora Architecture & Training Lifecycle

From raw tokenization to instruction tuning and low-rank parameter adaptation

Pipeline Flow
BPE Tokenizer IngestionTokenization
Vocab Size: 32,000

Transforms raw dialogues and markdown syntax into token IDs using custom Byte-Pair Encoding.

Transformer Forward PassModel Architecture
RMSNorm + RoPE + SwiGLU

Causal masking with RMSNorm pre-normalization, Rotary Position Embeddings (RoPE), and SwiGLU FFN.

Autoregressive PretrainingOptimization
Next-Token Loss

Pretrained on technical docs and repositories with AdamW optimizer and Cosine learning rate decay.

Supervised Fine-Tuning & LoRAAdaptation
Low-Rank Adaptation

Instruction fine-tuning using rank r=4 LoRA adapter matrices, training only ~1% of base parameters.

2. Architecture Variants: Micro vs 25M

NanoKora is implemented in PyTorch (kora/model.py) across two primary configurations:

NanoKora Micro
~902K Parameters (Ultra-compact)
  • • d_model: 128
  • • n_heads: 4
  • • n_layers: 4
  • • d_ff: 384
  • • max_seq_len: 256
NanoKora 25M
26.2M Parameters (Reasoning & Code)
  • • d_model: 512
  • • n_heads: 8
  • • n_layers: 8
  • • d_ff: 1408
  • • max_seq_len: 256

3. Transformer Components & Math

RMSNorm (Root Mean Square Normalization)

Omits mean-centering from standard LayerNorm, reducing GPU overhead by 20%:

kora/model.py (RMSNorm)
class RMSNorm(nn.Module):
    def __init__(self, dim: int, eps: float = 1e-6):
        super().__init__()
        self.eps = eps
        self.weight = nn.Parameter(torch.ones(dim))

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        variance = x.pow(2).mean(-1, keepdim=True)
        return x * torch.rsqrt(variance + self.eps) * self.weight
RoPE (Rotary Position Embeddings)

Rotates pairs of query/key coordinate vectors in 2D planes, enabling relative position encoding with natural decay across sequence distances.

SwiGLU Gated Feed-Forward Network

Uses SiLU-gated linear projections (matching LLaMA 3 and Mistral architectures) outperforming standard ReLU:

kora/model.py (SwiGLU)
class SwiGLUFFN(nn.Module):
    def __init__(self, dim: int, hidden_dim: int):
        super().__init__()
        self.w_gate = nn.Linear(dim, hidden_dim, bias=False)
        self.w_up = nn.Linear(dim, hidden_dim, bias=False)
        self.w_down = nn.Linear(hidden_dim, dim, bias=False)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        return self.w_down(F.silu(self.w_gate(x)) * self.w_up(x))

4. Tokenization & Vocabulary (BPE)

NanoKora uses a custom Byte-Pair Encoding (BPE) tokenizer (kora/tokenizer.py) supporting special dialogue delimiters:

Chat Template Format
<|im_start|>system
You are KORA, an expert software architect.<|im_end|>
<|im_start|>user
Refactor this Express route to Next.js App Router.<|im_end|>
<|im_start|>assistant

5. Parameter-Efficient LoRA Adaptation

Kora leverages Low-Rank Adaptation (kora/lora.py). Base attention weights (W_0) are frozen, and low-rank factor decomposition matrices A and B (rank r=4) are injected:

kora/lora.py
# Mathematical LoRA Forward Formulation:
# W_new = W_0 + (alpha / r) * (B @ A)

class LoRALinear(nn.Module):
    def __init__(self, linear: nn.Linear, r: int = 4, alpha: float = 8.0):
        super().__init__()
        self.linear = linear
        self.linear.weight.requires_grad = False  # Freeze base weights
        self.r = r
        self.scaling = alpha / r
        self.lora_A = nn.Parameter(torch.randn(r, linear.in_features) * 0.01)
        self.lora_B = nn.Parameter(torch.zeros(linear.out_features, r))

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        base_out = self.linear(x)
        lora_out = (x @ self.lora_A.T @ self.lora_B.T) * self.scaling
        return base_out + lora_out

6. Training Pipeline (Pretrain → SFT)

The training pipeline progresses in two discrete stages:

  • Stage 1: Causal Pretraining: Trained with next-token prediction cross-entropy loss, AdamW optimizer, and Cosine learning rate decay with linear warmup.
  • Stage 2: Instruction SFT: Trained exclusively on prompt-response pairs using response-only cross-entropy masking so prompt tokens do not penalize loss.

7. ONNX Export & Edge Inference

The model can be exported to ONNX format (kora/export_onnx.py) enabling client-side browser execution via WebAssembly / ONNX Runtime Web with zero cloud server compute costs.

Terminal
python kora/export_onnx.py --model_path checkpoints/nanokora_25m.pt --output nanokora_25m.onnx

8. OpenAI-Compatible Chat API

Query KORA directly through our standard API gateway endpoint:

cURL Request
curl -X POST https://api.kontyra.name.ng/v1/chat/completions \
  -H "Authorization: Bearer test-kontyra-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kora-v2",
    "messages": [
      { "role": "system", "content": "You are KORA, an expert software architect." },
      { "role": "user", "content": "Explain how SwiGLU outperforms ReLU." }
    ],
    "temperature": 0.7,
    "stream": false
  }'