MOI AI Architecture — MO Intelligence R&D Docs
Dense decoder-only transformer with RMSNorm, RoPE, GQA attention, SwiGLU feed-forward, and tied embeddings. R0: 100M params, d_model=768, 12 layers.
Dense decoder-only transformer with RMSNorm, RoPE, GQA attention, SwiGLU feed-forward, and tied embeddings. R0: 100M params, d_model=768, 12 layers.