Loading
Loading
Rigorous investigation of language model architectures, training methodology, fine-tuning techniques and evaluation — with published evidence and reproducible results.
Language models are at the centre of modern AI capability, but the field is dominated by a small number of large-scale actors whose training data, methodology and evaluation practices are often opaque. NOVACORE AI's language model research programme takes a different approach: we investigate what makes models effective on specific tasks and in specific languages — particularly Romanian — rather than pursuing indiscriminate scale. Our research focuses on architectures that deliver strong performance with fewer parameters, training methodologies that are transparent and reproducible, fine-tuning techniques that adapt models to domain-specific requirements without catastrophic forgetting, and evaluation frameworks that measure what matters for real-world deployment. Every research output includes published methodology, training-data provenance, evaluation results and model cards. We maintain a clear separation between research outputs (published with evidence) and commercial models (deployed for customers). The language model research programme proceeds through disciplined stage gates — from using established models through fine-tuning, continued pretraining and ultimately assessing whether proprietary foundation models are warranted by demand, data availability, staffing and compute capacity.
Investigating transformer variants, mixture-of-experts, state-space models and other architectures — with a focus on efficiency, inference speed and performance on non-English languages. Measuring the relationship between parameter count, training data volume and task performance.
Transparent training recipes — data composition, curriculum learning, optimisation schedules and hyperparameter selection. Documenting what works, what does not and why — with reproducible experiments and published training logs.
Techniques for domain adaptation, instruction tuning and task-specific optimisation. Parameter-efficient methods (LoRA, QLoRA, adapters), full fine-tuning trade-offs and evaluation of fine-tuned models against domain-specific benchmarks.
Comprehensive evaluation frameworks covering accuracy, robustness, fairness, toxicity and factual reliability. Task-specific benchmarks, few-shot evaluation protocols and human-evaluation methodologies with inter-annotator agreement measurements.
Research into model alignment — RLHF, DPO, constitutional AI and other approaches. Measuring the trade-off between helpfulness, harmlessness and honesty. Alignment techniques that preserve model capability while improving safety characteristics.
Systematic documentation of training data sources, licensing, filtering and composition. Tools and methodologies for auditing training data provenance — essential for regulatory compliance and research integrity.
| Track | Status | Output |
|---|---|---|
| Efficient architectures for Romanian | Active | Architecture comparison paper, open-weight models |
| Domain adaptation via fine-tuning | Active | Fine-tuning recipes, domain benchmarks |
| Training data documentation standards | Active | Data-provenance framework, tooling |
| Evaluation framework for non-English LLMs | In development | Benchmark suite, evaluation harness |
| Continued pretraining for Romanian | Planned | Continued-pretraining methodology, model release |
NOVACORE publishes language model research openly — methodology, evaluation results and model cards — but releases are governed. Models are published only when they meet documented quality, safety and documentation standards. Training data must be properly licensed and its provenance documented. Evaluation results must be reproducible. Research that involves commercial models or customer data follows separate governance — customer models and data are never used for public research without explicit, documented consent.
Secure AI and high-performance computing for enterprises, governments and research.