EvoLM-8B-Rubric

This repository contains the rubric-generator checkpoint from EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics. It is the final rubric checkpoint at global training step 1000 (rubric/step_1000) from the main EvoLM run. Because rubric and policy phases alternate, the rubric model has received 500 update steps at this point.

This is not the downstream policy checkpoint. The paired policy at global step 950 is published separately as stellalisy/EvoLM-8B.

Model details

  • Base model: Qwen/Qwen3-8B
  • Role: generate question-specific, weighted evaluation rubrics
  • Checkpoint: rubric/step_1000 from the main EvoLM run
  • Training data source: allenai/tulu-3-sft-mixture
  • Judge used in the main experiments: Qwen/Qwen3-1.7B

The model produces rubrics rather than scalar rewards. In the paper's evaluation and training pipeline, a separate judge model applies each generated rubric to candidate responses.

Usage

The checkpoint uses the standard Qwen chat template and can be loaded with Transformers:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "stellalisy/EvoLM-8B-Rubric"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

For paper-compatible generation, use the rubric_generation_v3 message format in open_instruct/search_rewards/utils/rubric_chat_templates.py. The system message asks for 2--5 atomic, weighted criteria in JSON; the user message is the question for which a rubric should be generated. Benchmark evaluation uses stella_run_scripts/paper_experiments/eval/eval_rewardbench2.sh and reward-bench/scripts/run_generative_v2_rubric.py.

Provenance

  • Artifact-relative source: outputs_v3_8ae3f0c/main/0319-e8ee435-v3_main_margin_format/v3_main_margin_format__1__1773964184_checkpoints/rubric/step_1000
  • Training run's recorded code revision: evolm-main-training-e8ee435 (e8ee4354c225c11faefb262bfd075d1a1216f1c5)
  • Nearest preserved working-tree snapshot: evolm-main-working-tree-e27c85c (e27c85c626f39649adbdbbf885684e0a21d14b1c)
  • Paper draft revision: a6234712f39aa145be2f4b72b45d7bbe6293b6e1

The training job recorded the first revision above but ran with uncommitted format-reward changes. The second revision is the nearest preserved snapshot containing the format-reward implementation and launcher observed in the run logs; it may also contain edits made shortly after launch.

Weight checksums

9c1aa0013e7fa75a738cf3a60d1f79fb97f4ea039570b624a35a270a09f013a4  model-00001-of-00004.safetensors
0dc89e4c869a250593768de5badb5c0892ffe48479745c3f2c6e5958a8e8edf2  model-00002-of-00004.safetensors
934e60a079d1f64cfd2313736a3f13d29a5f8c8c92ad29e2665aacb1129fd22b  model-00003-of-00004.safetensors
8cddce6f012c634bce61ad9b8c9e15529e55bb0ae10aa82347acb40668bc0c1c  model-00004-of-00004.safetensors

Limitations

This checkpoint was trained for rubric generation within the EvoLM pipeline. It is not a standalone answer-quality classifier, and rubric quality and final scores depend on the generation prompt, decoding settings, and judge model.

Downloads last month
278
Safetensors
Model size
308k params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for stellalisy/EvoLM-8B-Rubric

Finetuned
Qwen/Qwen3-8B
Finetuned
(2171)
this model
Quantizations
2 models