KORA (NanoKora Foundation Model)
The complete technical specification for Kora: modern decoder-only transformer architecture, RMSNorm, Rotary Position Embeddings (RoPE), SwiGLU feed-forward networks, and low-rank LoRA adaptation.
1. Overview & NanoKora Engine
KORA is Kontyra's native intelligence assistant, constructed from first principles to demystify modern large language model engineering while delivering fast, deterministic code assistance embedded directly within DevOS and Kontyra APIs.
From raw tokenization to instruction tuning and low-rank parameter adaptation
Transforms raw dialogues and markdown syntax into token IDs using custom Byte-Pair Encoding.
Causal masking with RMSNorm pre-normalization, Rotary Position Embeddings (RoPE), and SwiGLU FFN.
Pretrained on technical docs and repositories with AdamW optimizer and Cosine learning rate decay.
Instruction fine-tuning using rank r=4 LoRA adapter matrices, training only ~1% of base parameters.
2. Architecture Variants: Micro vs 25M
NanoKora is implemented in PyTorch (kora/model.py) across two primary configurations:
- • d_model: 128
- • n_heads: 4
- • n_layers: 4
- • d_ff: 384
- • max_seq_len: 256
- • d_model: 512
- • n_heads: 8
- • n_layers: 8
- • d_ff: 1408
- • max_seq_len: 256
3. Transformer Components & Math
Omits mean-centering from standard LayerNorm, reducing GPU overhead by 20%:
class RMSNorm(nn.Module):
def __init__(self, dim: int, eps: float = 1e-6):
super().__init__()
self.eps = eps
self.weight = nn.Parameter(torch.ones(dim))
def forward(self, x: torch.Tensor) -> torch.Tensor:
variance = x.pow(2).mean(-1, keepdim=True)
return x * torch.rsqrt(variance + self.eps) * self.weightRotates pairs of query/key coordinate vectors in 2D planes, enabling relative position encoding with natural decay across sequence distances.
Uses SiLU-gated linear projections (matching LLaMA 3 and Mistral architectures) outperforming standard ReLU:
class SwiGLUFFN(nn.Module):
def __init__(self, dim: int, hidden_dim: int):
super().__init__()
self.w_gate = nn.Linear(dim, hidden_dim, bias=False)
self.w_up = nn.Linear(dim, hidden_dim, bias=False)
self.w_down = nn.Linear(hidden_dim, dim, bias=False)
def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.w_down(F.silu(self.w_gate(x)) * self.w_up(x))4. Tokenization & Vocabulary (BPE)
NanoKora uses a custom Byte-Pair Encoding (BPE) tokenizer (kora/tokenizer.py) supporting special dialogue delimiters:
<|im_start|>system
You are KORA, an expert software architect.<|im_end|>
<|im_start|>user
Refactor this Express route to Next.js App Router.<|im_end|>
<|im_start|>assistant5. Parameter-Efficient LoRA Adaptation
Kora leverages Low-Rank Adaptation (kora/lora.py). Base attention weights (W_0) are frozen, and low-rank factor decomposition matrices A and B (rank r=4) are injected:
# Mathematical LoRA Forward Formulation:
# W_new = W_0 + (alpha / r) * (B @ A)
class LoRALinear(nn.Module):
def __init__(self, linear: nn.Linear, r: int = 4, alpha: float = 8.0):
super().__init__()
self.linear = linear
self.linear.weight.requires_grad = False # Freeze base weights
self.r = r
self.scaling = alpha / r
self.lora_A = nn.Parameter(torch.randn(r, linear.in_features) * 0.01)
self.lora_B = nn.Parameter(torch.zeros(linear.out_features, r))
def forward(self, x: torch.Tensor) -> torch.Tensor:
base_out = self.linear(x)
lora_out = (x @ self.lora_A.T @ self.lora_B.T) * self.scaling
return base_out + lora_out6. Training Pipeline (Pretrain → SFT)
The training pipeline progresses in two discrete stages:
- Stage 1: Causal Pretraining: Trained with next-token prediction cross-entropy loss, AdamW optimizer, and Cosine learning rate decay with linear warmup.
- Stage 2: Instruction SFT: Trained exclusively on prompt-response pairs using response-only cross-entropy masking so prompt tokens do not penalize loss.
7. ONNX Export & Edge Inference
The model can be exported to ONNX format (kora/export_onnx.py) enabling client-side browser execution via WebAssembly / ONNX Runtime Web with zero cloud server compute costs.
python kora/export_onnx.py --model_path checkpoints/nanokora_25m.pt --output nanokora_25m.onnx8. OpenAI-Compatible Chat API
Query KORA directly through our standard API gateway endpoint:
curl -X POST https://api.kontyra.name.ng/v1/chat/completions \
-H "Authorization: Bearer test-kontyra-key" \
-H "Content-Type: application/json" \
-d '{
"model": "kora-v2",
"messages": [
{ "role": "system", "content": "You are KORA, an expert software architect." },
{ "role": "user", "content": "Explain how SwiGLU outperforms ReLU." }
],
"temperature": 0.7,
"stream": false
}'