Skip to main content
1 min readKnowledge Resource

Knowledge Resource · Open access

Research Summary: Greenpixie's AI Token Methodology: Assessing the Energy, Water and CO2-eq Impact of AI Tokens for Open and Closed Weight Models

Original authors
Attribution requires verification
Original source
arXiv — Computers and Society
Summary & Analysis prepared by
Aziz Shuaib Ausi
Resource type
Research Summary / Knowledge Resource
Resource published on AZIZ OS
3 October 2026
Reading time
1 min
Publication type
Knowledge Resource
Availability
Open access
About this Summary & Analysis

AZIZ OS provides independently prepared summaries and analytical interpretations of externally published research and knowledge sources. The underlying works remain attributable to their original authors and rights holders. This resource is intended to improve accessibility and understanding and does not replace the original publication.

Checking access…

The research describes a novel methodology for assessing the per-token energy, water, and CO2-equivalent impact of cloud-hosted Large Language Model (LLM) inference. This approach differentiates between input (prefill) and output (decode) tokens, primarily focusing on Graphics Processing Unit (GPU) energy usage during inference benchmarking. It incorporates a model that relates energy per token to LLM size, traffic, and hardware configuration, extending its application to proprietary models through size binning.

Why it matters

This methodology provides a standardized and granular approach to quantify the environmental footprint of AI model usage, offering critical data for sustainability initiatives and resource optimization. Understanding the energy costs at a per-token level enables more informed decision-making regarding AI deployment, hardware selection, and operational efficiency across various sectors reliant on LLMs.

Key insights

  • A methodology has been developed to estimate the per-token energy cost for cloud-hosted LLM inference.
  • The methodology distinguishes between energy costs for input (prefill) and output (decode) tokens.
  • GPU energy consumption is directly measured during inference benchmarking for open-weight models across various text-based tasks.
  • Non-GPU server energy contributions are estimated based on inference wall time.
  • Bayesian linear regression is employed to model the relationship between energy per token and LLM size, request traffic, and hardware deployment.
  • Proprietary frontier LLMs are categorized into size buckets using naming conventions and performance indicators for impact assessment.

Source

arXiv — Computers and Society — https://arxiv.org/abs/2609.33965

Citation

Cite the original work (APA 7)

The original source is authoritative for this citation. Cite the source publication directly — this attribution is pending verification. Open the original source.

Verification

This is an authenticated AZIZ OS resource record.

Verification ID
ASA-EXE-2026-01144
Version
v1.0 · r0
Issued
3 October 2026
Resource prepared by
Aziz Shuaib Ausi
Resource status
Research Summary / Knowledge Resource
Underlying work
Greenpixie's AI Token Methodology: Assessing the Energy, Water and CO2-eq Impact of AI Tokens for Open and Closed Weight Models
Original authors
Attribution requires verification
Original source
arXiv — Computers and Society
Provenance status
Attribution requires verification
Rights
Underlying publication rights remain with the respective copyright holder(s). Refer to the original source for authoritative publication and licensing information.

This verification confirms the AZIZ OS resource record and its documented provenance. It does not establish authorship of the underlying external work.

Verify this resource