Research Paper · Information Technology

GreenServe: Energy-Aware Scheduling of Large Language Model Inference Workloads in Serverless Edge Computing Environments

An original research paper proposing a carbon-aware, multi-tier scheduling framework for LLM inference. Reported results: 31.4% energy reduction and 42.7% carbon reduction versus a latency-only baseline, with 97.8% SLO attainment.

Download

Paper structure

  1. Abstract & Keywords
  2. 1. Introduction (motivation, problem statement, RQ1–RQ4, contributions)
  3. 2. Literature Review (18 peer-reviewed sources)
  4. 3. System Model
  5. 4. GreenServe Design (cost model, carbon-aware placement, pre-warming)
  6. 5. Experimental Methodology
  7. 6. Results (energy, carbon, latency, ablations, sensitivity)
  8. 7. Discussion (implications, embodied carbon, threats to validity, ethics)
  9. 8. Conclusion & Future Work
  10. References, Appendix A (notation), Appendix B (reproducibility)