Boosted by baldur@toot.cafe ("Baldur Bjarnason"):
jcoglan wrote:
somehow I only recently learned that LLMs need the whole conversation fed back to them on each prompt so their i/o cost scales as O(n^2), something that would be considered completely unacceptable in almost any other production network-accessible software