Published October 6, 2025 | Version v1.0.0

nafew-azim/AUTODSPy: v1.0.0 – Initial Stable Release

Authors/Creators

Description

AutoDSPy: Automating Modular Prompt Design with Reinforcement Learning

AutoDSPy is a groundbreaking framework that fully automates the construction of DSPy pipelines for small and large language models (LLMs) using reinforcement learning (RL). By leveraging an RL-tuned policy network, it dynamically selects optimal reasoning modules (e.g., Chain-of-Thought for logical tasks, ReAct for tool integration), input-output signatures, and execution strategies—eliminating manual configuration and making LLMs more accessible and scalable.

As presented in our EMNLP 2025 paper: AutoDSPy: Automating Modular Prompt Design with Reinforcement Learning for Small and Large Language Models.

Highlights from Initial Release (v1.0.0)

  • Empirical Gains: Achieves up to 4.3% accuracy improvement over DSPy baselines on GSM8K (mathematical reasoning) and HotPotQA (multi-hop QA), while reducing inference time—even with compact models like GPT-2 (127M parameters).
  • RL Strategies: Implements three RL algorithms—REINFORCE (with baseline), PPO (Proximal Policy Optimization), and GRPO (Group Relative Policy Optimization)—for training a policy network to generate task-adaptive pipelines.
  • Modular and Efficient: Supports predefined modules (Predict, CoT, ReAct), 15+ signatures (e.g., "question -> answer"), and zero-shot or few-shot optimization without user intervention.
  • Use Cases: Ideal for complex reasoning tasks like math problem-solving, multi-hop QA, and beyond—extensible to code synthesis, scientific reasoning, and multimodal applications.

Quick Start

Prerequisites

  • Python 3.8+
  • Install dependencies: pip install torch transformers dspy-ai datasets sentence-transformers numpy ollama
  • Start Ollama server: ollama serve and pull a model (e.g., ollama pull llama3.1:8b)

Training a Model

Run one of the provided scripts:

python reinforce_training.py  # Simple REINFORCE with baseline
python ppo_training.py       # PPO for stable optimization
python grpo_training.py      # GRPO for group-relative advantages

Models are saved automatically (e.g., gpt2_trained_policy_model.pt).

Testing

Load and test a trained model:

import torch
from transformers import GPT2LMHeadModel, GPT2Tokenizer

policy_model = GPT2LMHeadModel.from_pretrained('gpt2')
policy_model.load_state_dict(torch.load('gpt2_trained_policy_model.pt'))
policy_model.eval()

# Test on a prompt
test_model(policy_model, "Solve: Stacy has 36 oranges and ends with 20. How many sold?")

Evaluation

Evaluate on benchmarks:

evaluate_module(gsm8k_test_data, gsm8k_test, policy_model)  # For GSM8K
evaluate_module(hotpotqa_test_data, hotpotqa_test, policy_model, query=True)  # For HotPotQA

For full details, see the README.md for step-by-step guides, troubleshooting, and customization tips.

Contributions

  • Framework: First fully automated DSPy extension via RL.
  • Policy Network: RL-tuned LLM for dynamic pipeline synthesis.
  • Benchmarks: Superior accuracy/efficiency on GSM8K and HotPotQA.

Star the repo if you find it useful! Contributions welcome—fork and PR for enhancements.

Developed by the Apurba-NSU R&D Lab team.

There was an error creating your Release: tag name can't be blank, tag name is not well-formed, published releases must have a valid tag.

Files

nafew-azim/AUTODSPy-v1.0.0.zip

Files (24.8 kB)

Name Size Download all
md5:87f7abe43f689197f0d07b5ce25f1ff8
24.8 kB Preview Download

Additional details

Related works