Research Paper · Information Technology
GreenServe: Energy-Aware Scheduling of Large Language Model Inference Workloads in Serverless Edge Computing Environments
An original research paper proposing a carbon-aware, multi-tier scheduling framework for LLM inference. Reported results: 31.4% energy reduction and 42.7% carbon reduction versus a latency-only baseline, with 97.8% SLO attainment.
Download
Paper structure
- Abstract & Keywords
- 1. Introduction (motivation, problem statement, RQ1–RQ4, contributions)
- 2. Literature Review (18 peer-reviewed sources)
- 3. System Model
- 4. GreenServe Design (cost model, carbon-aware placement, pre-warming)
- 5. Experimental Methodology
- 6. Results (energy, carbon, latency, ablations, sensitivity)
- 7. Discussion (implications, embodied carbon, threats to validity, ethics)
- 8. Conclusion & Future Work
- References, Appendix A (notation), Appendix B (reproducibility)