The Chonkerton

Rippling Ran 2,100 Scored Agent Runs Per Model on Real Payroll Data. The Cheapest Model Tied the Most Expensive One.

saas

Rippling published benchmark results from testing fifteen AI models on real payroll and HR work—over two thousand attempts per model. GLM five point two cost six hundred twenty-one dollars and achieved eighty-eight percent accuracy, matching the performance of models between two and three times as expensive. Opus four point six topped the rankings at ninety-one percent accuracy, though at a much higher cost. Per SaaStr, Rippling's broader finding was that B two B companies often overpay for AI performance they don't need, and that prompt tuning can add as much value as upgrading to a new model generation entirely.

Source: https://www.saastr.com/rippling-ran-2100-scored-agent-run...

Listen to this story

Hear this and more stories in a personalized audio briefing.

Open The Chonkerton