Tor to the World · story 1138 · unverified · 1 source(s)

42x faster prompt lookup drafting in llama.cpp

A new optimization in llama.cpp achieves 42 times faster prompt lookup drafting.

Open in the desk

Coverage

What this site indexes