Small local models go far
Small models in the browser with WebGPU go far for FAQs with good retrieval. Total privacy, zero server cost.
The secret is in the chunks and the prompt, not the parameters. A 360M model with the right three paragraphs beats a giant one guessing.
This site's chat widget is the experiment running live.