Hugging Face
Models
Datasets
Spaces
Posts
Docs
Enterprise
Pricing
Log In
Sign Up
CohenQu
/
DeepSeek-R1-Distill-Qwen-7B-GRPO
like
4
Text Generation
Transformers
Safetensors
hf-cmu-collab/DeepScaleR-1.5B-Preview_on-policy_GRPO
qwen2
Generated from Trainer
trl
grpo
conversational
text-generation-inference
Inference Endpoints
arxiv:
2402.03300
Model card
Files
Files and versions
Community
Train
Deploy
Use this model
main
DeepSeek-R1-Distill-Qwen-7B-GRPO
/
config.json
Commit History
Training in progress, step 1
5ac6ce6
verified
CohenQu
commited on
4 days ago
Training in progress, step 1
054fce6
verified
CohenQu
commited on
4 days ago
Training in progress, step 1
b368e6f
verified
CohenQu
commited on
6 days ago
Training in progress, step 12
1556ad5
verified
CohenQu
commited on
8 days ago
Training in progress, step 10
8349a49
verified
CohenQu
commited on
8 days ago
Training in progress, step 10
dfef9e8
verified
CohenQu
commited on
8 days ago
Training in progress, step 8
4c7ac98
verified
CohenQu
commited on
8 days ago
Training in progress, step 8
aee2252
verified
CohenQu
commited on
8 days ago
Training in progress, step 6
be41817
verified
CohenQu
commited on
8 days ago
Training in progress, step 6
93a1c82
verified
CohenQu
commited on
8 days ago
Training in progress, step 4
8adb8de
verified
CohenQu
commited on
8 days ago
Training in progress, step 4
c5c2905
verified
CohenQu
commited on
8 days ago
Training in progress, step 2
77b71d6
verified
CohenQu
commited on
8 days ago
Training in progress, step 2
28a68fb
verified
CohenQu
commited on
8 days ago
Training in progress, step 1
c961acf
verified
CohenQu
commited on
8 days ago
Training in progress, step 4
45b7fce
verified
CohenQu
commited on
9 days ago
Training in progress, step 1
cc2ac4d
verified
CohenQu
commited on
13 days ago
Training in progress, step 20
17e8d9c
verified
CohenQu
commited on
15 days ago