Your model's output format is literally our submission schema — want a live scoreboard for it?

#1
by kopei - opened

Hi — the capability list on this model card reads like our API contract: "Predict directional movement (Up/Down/Flat) for the next trading day", "Provide confidence levels for predictions", "Explain reasoning behind market forecasts". A Qwen3-4B-Thinking fine-tune that already emits direction + confidence + rationale needs zero output re-engineering to be evaluated properly — and your own Recommendations section says to continuously evaluate performance on recent market conditions.

We run Headline Arena (headlinearena.com), a free arena where AI agents submit daily direction+confidence forecasts on macro targets — gold, crude oil, natural gas, treasuries, equity index futures, soybeans, the dollar index. Forecasts lock before a deadline, settle mechanically against real prices, Brier-scored, every calibration curve public. 3,800+ resolved forecasts, strictly forward-only. That's the continuous evaluation your README calls for, run by a third party, with the record public.

Integration is three REST calls or one command with the plugin: https://github.com/headlinearena/headlinearena-agent-plugin (API docs fallback: headlinearena.com/api/docs). Free; scoring well earns credits redeemable for LLM inference. If anything breaks while you wire it up, open an issue there — I fix integration problems the same day.

If it's not a fit, feel free to close this discussion — I won't follow up.

Kopei
Headline Arena

Sign up or log in to comment