Loading
Loading
Building genuine AI capability for the Romanian language — datasets, evaluation frameworks, fine-tuned models and benchmarks developed in Romania, for Romanian speakers.
Romanian is spoken by approximately 24 million people, making it one of the larger European languages — yet it is severely underrepresented in AI. Most large language models are developed for English, with other languages receiving secondary attention at best. When Romanian-language support exists, it is typically the result of incidental inclusion in multilingual training corpora rather than deliberate engineering for the specific characteristics of the Romanian language — its rich morphology, case system, diacritics, regional variations and the particular challenges of processing Romanian legal, medical and administrative text. NOVACORE AI is committed to building genuine Romanian-language AI capability. This means more than translating English models or including Romanian text in training data. It means curating high-quality Romanian datasets, developing evaluation frameworks that measure real Romanian-language performance (not just translated English benchmarks), fine-tuning models specifically for Romanian tasks and publishing evidence about what works — and what does not — for this specific language. As a Romanian company operating in Europe, this is both a commercial priority and a matter of digital sovereignty: Romanian speakers deserve AI systems that work as well in their language as in English, and Romanian institutions deserve AI infrastructure that understands Romanian law, Romanian administrative practice and Romanian cultural context.
Curating high-quality Romanian-language datasets across domains — legal, medical, administrative, literary, journalistic and technical. All datasets are properly licensed, documented with provenance metadata and released with clear usage terms. Building the data foundation that Romanian AI has been missing.
Fine-tuned and continued-pretraining models optimised for Romanian. Models that understand Romanian morphology, handle diacritics correctly, respect formal and informal registers, and perform well on Romanian-specific tasks — not just translated English benchmarks.
Romanian-language evaluation benchmarks developed by native speakers — not translated English tests. Measuring grammar, diacritic usage, register appropriateness, legal and administrative comprehension, and cultural-context understanding. Published with methodology and results.
Applied research demonstrating Romanian AI in real-world contexts — government document processing, legal-text analysis, medical-record summarisation, educational content generation and citizen-service automation. Evidence of what Romanian AI can actually deliver.
Developing and releasing linguistic resources for Romanian NLP — annotated corpora, dependency treebanks, named-entity recognition datasets, sentiment lexicons and morphological analysers. Building the infrastructure that enables others to build Romanian AI.
Partnerships with Romanian universities, research institutes, government agencies and cultural institutions. Joint dataset development, shared evaluation frameworks and co-authored research publications. Romanian AI built with Romanian institutions.
Systematic evaluation of how well current models perform on Romanian-language tasks. Publish honest benchmarks showing where Romanian AI stands today — and where the biggest gaps are.
Assemble, license, clean and document high-quality Romanian text across domains. Release datasets with clear provenance, usage terms and quality metrics.
Domain-specific fine-tuning of established models on Romanian data. Publish fine-tuning recipes, evaluation results and model weights.
Create comprehensive evaluation frameworks designed by Romanian linguists and domain experts. Benchmarks that measure what matters for Romanian — not just translated English tests.
Applied research demonstrating Romanian AI in government, healthcare, legal and educational contexts. Reference implementations with published performance evidence.
AI that does not work well in Romanian is not neutral — it disadvantages Romanian speakers, Romanian institutions and the Romanian economy. Government agencies cannot process citizen documents efficiently. Healthcare providers cannot summarise patient records accurately. Legal professionals cannot analyse Romanian case law reliably. Students cannot learn in their native language. Building genuine Romanian-language AI is not a cultural nicety — it is an economic necessity and a matter of digital sovereignty. NOVACORE AI is committed to closing the Romanian AI gap with properly resourced, evidence-based research — published openly, developed collaboratively and deployed responsibly.
Secure AI and high-performance computing for enterprises, governments and research.