Antonio V. Franco

My work focuses on AI research and inference engineering for LLMs and SLMs, with an emphasis on quantization, fine-tuning, low-precision computation, and GPU kernels in Triton/CUDA.

For partnerships and projects, contact me by email: contact@antoniovfranco.com

You can read more about me on


Articles