Model introduction
Leaderboard Submission
Leaderboard submission for a custom RL imitated expert model "NemoSmith-4B-Instruct" built on Nemotron-4B-Instruct for GPU compiler and assembly specific code generation.
Model & Framework
- Model: NemoSmith-4B-Instruct (base: nvidia/Nemotron-Mini-4B-Instruct + RL imitation for compilers)
- Size: 4B (open weights)
- Inference mode: Greedy (temperature=0), single attempt, chat setting
Results
| Benchmark |
pass@1 |
| HumanEval (base) |
35.4% |
| HumanEval+ (base + extra) |
32.9% |
| MBPP (base) |
43.4% |
| MBPP+ |
36.5% |
Reproducibility
Evaluated with the official evalplus.evaluate (v0.3.1), greedy decoding, served via vLLM (OpenAI backend).
Model URL
https://huggingface.co/datasets/abhilash1910/nemosmith-evalplus-samples
Additional information (Optional)
No response
Decontamination
CoT Inference refinement of Claude /Codex with custom RL imitation for compilation tasks and assembly languages (GPU) .
Author
Yes
Data
No
Security
- I confirm that the model is safe to run which is not designed to produce malicious code or content.
Integrity
- I confirm that the model comes from unique and original work and does not contain any plagiarism.
@ganler for a review.
Reactions are currently unavailable
Model introduction
Leaderboard Submission
Leaderboard submission for a custom RL imitated expert model "NemoSmith-4B-Instruct" built on Nemotron-4B-Instruct for GPU compiler and assembly specific code generation.
Model & Framework
Results
Reproducibility
Evaluated with the official evalplus.evaluate (v0.3.1), greedy decoding, served via vLLM (OpenAI backend).
Model URL
https://huggingface.co/datasets/abhilash1910/nemosmith-evalplus-samples
Additional information (Optional)
No response
Decontamination
CoT Inference refinement of Claude /Codex with custom RL imitation for compilation tasks and assembly languages (GPU) .
Author
Yes
Data
No
Security
Integrity
@ganler for a review.