Evaluate on your own data
Score Grayson on your own past cases with the grayson CLI. Point it at a CSV, ask your questions, and see how often Grayson gives the answer you know is right.
The grayson command-line tool sends your past cases to Grayson and compares its answers with the answers you know are right. You get the accuracy for each question, and a CSV with Grayson's answer next to yours for every case.
Install
pipx install https://docs.finic.ai/downloads/grayson_cli-0.2.2-py3-none-any.whlIt needs Python 3.9 or later, and an API key from portal.finic.ai/keys.
Prepare a CSV
One row per case, with a column that holds the right answer for each question you'll ask. Every other column is sent to Grayson as the case. An id column, if you have one, names the rows in the results.
id,account_age_days,amount_usd,merchant,confirmed_fraud
1042,19,2400,Electronics store,yes
1043,1260,58,Grocery store,noUse cases whose outcome you know, with only what was known when the decision was made. A later chargeback or a fraud label in the other columns makes Grayson look better than it will be.
Run it
graysonIt asks for:
- Your CSV.
- Each question: its type (yes/no, choice or score), its wording, and the column with the right answer: type to filter the columns and use the arrow keys to pick. A choice takes its options from the answers in that column; a score uses the standard likelihood levels or yours.
- Your API key.
Then it runs the cases at 5 per second, the standard rate limit, and shows each question's accuracy. 1,000 cases of about 1,000 tokens each cost about $0.04 and take about 3 minutes. Stop at any point and run it again to pick up where it left off.
What you get
A folder named after your file, such as cases.grayson/, with:
results.csv: one row per case. For each question, your answer, Grayson's, its probability and whether they match, then the latency and tokens.report.md: for each question, accuracy, precision and recall, and a table for choosing the probability at which to act.disagreements.csv: the cases where Grayson's answer differs from yours, most confident first. Check these first: they are often wrong labels.
Your cases go only to the Grayson API. Whether Finic keeps them follows your organization's data retention setting.
Run it again without the questions
At the end, grayson prints the command that repeats the run, with your questions saved as questions.json:
grayson eval cases.csv --questions cases.grayson/questions.json --label is_this_account_fraud=confirmed_fraud| Flag | Does |
|---|---|
--questions | The questions, as JSON: a file, or a URL such as a recipe's questions.json. These can also be multiple choice |
--label question=column | The column with that question's right answer. Repeat for each question |
--threshold question=p | Count a yes/no answer as yes from probability p (default 0.5) |
--limit n | Only the first n cases |
--dry-run | Check the file and estimate the cost; send nothing |
-y | Don't ask before sending |
To compare Grayson with your current rules or model, add a baseline_<question> column with what it answered; the report puts the two side by side.