nafew-azim/AUTODSPy: v1.0.0 – Initial Stable Release
Authors/Creators
Description
AutoDSPy: Automating Modular Prompt Design with Reinforcement Learning
AutoDSPy is a groundbreaking framework that fully automates the construction of DSPy pipelines for small and large language models (LLMs) using reinforcement learning (RL). By leveraging an RL-tuned policy network, it dynamically selects optimal reasoning modules (e.g., Chain-of-Thought for logical tasks, ReAct for tool integration), input-output signatures, and execution strategies—eliminating manual configuration and making LLMs more accessible and scalable.
As presented in our EMNLP 2025 paper: AutoDSPy: Automating Modular Prompt Design with Reinforcement Learning for Small and Large Language Models.
Highlights from Initial Release (v1.0.0)
- Empirical Gains: Achieves up to 4.3% accuracy improvement over DSPy baselines on GSM8K (mathematical reasoning) and HotPotQA (multi-hop QA), while reducing inference time—even with compact models like GPT-2 (127M parameters).
- RL Strategies: Implements three RL algorithms—REINFORCE (with baseline), PPO (Proximal Policy Optimization), and GRPO (Group Relative Policy Optimization)—for training a policy network to generate task-adaptive pipelines.
- Modular and Efficient: Supports predefined modules (Predict, CoT, ReAct), 15+ signatures (e.g., "question -> answer"), and zero-shot or few-shot optimization without user intervention.
- Use Cases: Ideal for complex reasoning tasks like math problem-solving, multi-hop QA, and beyond—extensible to code synthesis, scientific reasoning, and multimodal applications.
Quick Start
Prerequisites
- Python 3.8+
- Install dependencies:
pip install torch transformers dspy-ai datasets sentence-transformers numpy ollama - Start Ollama server:
ollama serveand pull a model (e.g.,ollama pull llama3.1:8b)
Training a Model
Run one of the provided scripts:
python reinforce_training.py # Simple REINFORCE with baseline
python ppo_training.py # PPO for stable optimization
python grpo_training.py # GRPO for group-relative advantages
Models are saved automatically (e.g., gpt2_trained_policy_model.pt).
Testing
Load and test a trained model:
import torch
from transformers import GPT2LMHeadModel, GPT2Tokenizer
policy_model = GPT2LMHeadModel.from_pretrained('gpt2')
policy_model.load_state_dict(torch.load('gpt2_trained_policy_model.pt'))
policy_model.eval()
# Test on a prompt
test_model(policy_model, "Solve: Stacy has 36 oranges and ends with 20. How many sold?")
Evaluation
Evaluate on benchmarks:
evaluate_module(gsm8k_test_data, gsm8k_test, policy_model) # For GSM8K
evaluate_module(hotpotqa_test_data, hotpotqa_test, policy_model, query=True) # For HotPotQA
For full details, see the README.md for step-by-step guides, troubleshooting, and customization tips.
Contributions
- Framework: First fully automated DSPy extension via RL.
- Policy Network: RL-tuned LLM for dynamic pipeline synthesis.
- Benchmarks: Superior accuracy/efficiency on GSM8K and HotPotQA.
Star the repo if you find it useful! Contributions welcome—fork and PR for enhancements.
Developed by the Apurba-NSU R&D Lab team.
There was an error creating your Release: tag name can't be blank, tag name is not well-formed, published releases must have a valid tag.
Files
nafew-azim/AUTODSPy-v1.0.0.zip
Files
(24.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:87f7abe43f689197f0d07b5ce25f1ff8
|
24.8 kB | Preview Download |
Additional details
Related works
- Is supplement to
- Software: https://github.com/nafew-azim/AUTODSPy/tree/v1.0.0 (URL)
Software
- Repository URL
- https://github.com/nafew-azim/AUTODSPy