Hello Claire, thank you for the explanation. But I think you probably not understand my questions.
For your points 1–9, I think they can all be implemented in user space. Yes, translators can bring these benefits. But my question is: Why does it need to be implemented as a translator rather than as a program? If you want to modularize the LLM, I think you can think about how to handle the APIs between the programs. For your point 10, I think the efficiency of LLMs today mainly comes from using GPUs, with only some edge cases using CPUs. If you use a CPU, only relatively small models, such as 9B models or some small MLP/CNN/PINN, can run reasonably well. For your points 11 and 12, I think they are the same as my points above: these things can also be implemented by ordinary user-space programs. Yes, translators can provide these properties, but why translators specifically?
