Hi William, 
I too would have liked to learn more about this but here’s a simple explanation that I gleaned from my limited knowledge and chatGTP. 
As you are aware the many  AI Large language models (LLMs) that underly chatGTP, Gemini, Claude, DeepSeek etc are trained by ‘supplying’ them with huge amounts of training data like magazines, journals and who know what! This is a highly contentious issue with creators which is why NYT has sued them but that is another story. 
But these LLMs  struggle with languages like Konkani because there isn’t enough training data and the language is written in multiple scripts (like Devanagari, Romi, and Kannada), which makes learning (for these models) harder. To fix this, the authors created a large synthetic dataset  (Konkani-Instruct-100k) and used it to fine-tune existing models such as Llama 3.1 (meta) Qwen2.5 (Ali baba) and Gemma 3 (Google). They also built benchmarks to evaluate performance across scripts. As a result, their improved models perform better at tasks like translation, sometimes even outperforming propriety baselines. 
As Joǎo said if you want to know how they do it then you need expertise in computer sciences and math but as a highly regarded Konkani researcher I think you would be a perfect user to test their fine tuned model.  
Hope this helps.
Best, 
Helga 
Sent from my iPhone

On Apr 10, 2026, at 06:22, William Robert Da Silva <[email protected]> wrote:


Send some clarification on the entire project.
WRDS

On Mon, Apr 6, 2026 at 12:49 AM Frederick Noronha <[email protected]> wrote:

Konkani LLM: Multi-Script Instruction Tuning and Evaluation for a Low-Resource Indian Language

Large Language Models (LLMs) consistently under perform in low-resource linguistic contexts such as Konkani. This performance deficit stems from acute training data scarcity compounded by high script diversity across Devanagari, Romi and Kannada orthographies. To address this gap, we introduce Konkani-Instruct-100k, a comprehensive synthetic instruction-tuning dataset generated through Gemini 3.
We establish rigorous baseline benchmarks by evaluating leading open-weights architectures including Llama 3.1, Qwen2.5 and Gemma 3 alongside proprietary closed-source models. Our primary contribution involves the development of Konkani LLM, a series of fine-tuned models optimized for regional nuances. Furthermore, we are developing the Multi-Script Konkani Benchmark to facilitate cross-script linguistic evaluation. In machine translation, Konkani LLM delivers consistent gains over the corresponding base models and is competitive with and in several settings surpasses proprietary baselines  https://arxiv.org/abs/2603.23529


_/_/_/_/_/_/_/_/_/_/_/_/_/_/_/_/_/
_/  Frederick Noronha  फ्रेडरिक नोरोन्या  * فريدريك نورونيا‎
_/  AUDIO https://archive.org/details/@fredericknoronha
_/  http://goa1556.in +91-9822122436 784 Saligao Goa
_/_/_/_/_/_/_/_/_/_/_/_/_/_/_/_/_/

--
You received this message because you are subscribed to the Google Groups "Goa-Research-Net" group.
To unsubscribe from this group and stop receiving emails from it, send an email to [email protected].
To view this discussion, visit https://groups.google.com/d/msgid/goa-research-net/CA%2Bmqab8n87oQUuD51v0UMO5e2-o7Kb-cehxiZz7q7JePCndJVA%40mail.gmail.com.

--
You received this message because you are subscribed to the Google Groups "Goa-Research-Net" group.
To unsubscribe from this group and stop receiving emails from it, send an email to [email protected].
To view this discussion, visit https://groups.google.com/d/msgid/goa-research-net/CA%2BvNr4%2BU_ms-95dM9C4UNcMj-6tE4-nAyW8aPtGnC_9_5dJmFQ%40mail.gmail.com.

--
You received this message because you are subscribed to the Google Groups "Goa-Research-Net" group.
To unsubscribe from this group and stop receiving emails from it, send an email to [email protected].
To view this discussion, visit https://groups.google.com/d/msgid/goa-research-net/9F54C4CB-0CBC-4A67-86FC-CBC56A86883A%40gmail.com.

Reply via email to