(github.com)

268 points Areibman | 1 comments | 17 Jun 24 19:44 UTC | HN request time: 0.193s | source

Hey HN! Tokencost is a utility library for estimating LLM costs. There are hundreds of different models now, and they all have their own pricing schemes. It’s difficult to keep up with the pricing changes, and it’s even more difficult to estimate how much your prompts and completions will cost until you see the bill.

Tokencost works by counting the number of tokens in prompt and completion messages and multiplying that number by the corresponding model cost. Under the hood, it’s really just a simple cost dictionary and some utility functions for getting the prices right. It also accounts for different tokenizers and float precision errors.

Surprisingly, most model providers don't actually report how much you spend until your bills arrive. We built Tokencost internally at AgentOps to help users track agent spend, and we decided to open source it to help developers avoid nasty bills.

Show context

Lerc ◴[17 Jun 24 21:36 UTC] No.40711341[source]▶

>>40710154 (OP) #

With all the options there seems like an opportunity for a single point API that can take a series of prompts, a budget and a quality hint to distribute batches for most bang for buck.

Maybe a small triage AI to decide how effectively models handle certain prompts to preserve spending for the difficult tasks.

Does anything like this exist yet?

replies(3): >>40712921 #>>40715521 #>>40715879 #

curious_cat_163 ◴[18 Jun 24 00:47 UTC] No.40712921[source]▶

>>40711341 #

I have yet to find a use case where quality can be traded off.

Would love to hear what you had in mind.

replies(3): >>40713311 #>>40713472 #>>40788282 #

1. Lerc ◴[18 Jun 24 02:19 UTC] No.40713472[source]▶

>>40712921 #

It is not so much a drop in quality as there are tasks that every model above a certain threshold will perform equally.

Most can do 2+2 = 4.

One test prompt I use on LLMs is asking it to produce a JavaScript function that takes an ImageData object and returns a new ImageData object with an all direction Sobel edge detection. Quite a lot of even quite small models can generate functions like this.

In general, I don't even think this is a question that needs to be answered. A lot of API providers have different quality/price tiers. The fact that people are using the different tiers should be sufficient to show that at least some people are finding cases where cheaper models are good enough.

↑

Show HN: Token price calculator for 400+ LLMs