How to Authenticate and Fetch Data from Google Analytics 4 and Search Console in Python

Both modules authenticate via service-account JSON keys stored in environment variables, then fetch data using the BetaAnalyticsDataClient for GA4 and the googleapiclient discovery service for Search Console.

The seomachine repository provides Python wrappers that simplify extracting SEO metrics from Google’s APIs. Understanding how google_analytics.py and google_search_console.py handle authentication and data retrieval lets you integrate these data sources securely into your own pipelines without hardcoding secrets.

Service-Account Authentication Architecture

Both modules follow an identical service-account pattern: they load a Google Cloud JSON key from disk, create scoped credentials, and initialize a client object that persists for the lifetime of the class instance.

Environment Variable Configuration

Secrets are externalized to environment variables to keep the codebase free of credentials. The modules read the following keys during initialization:

Credential Loading and Client Initialization

In data_sources/modules/google_analytics.py, the GoogleAnalytics.__init__ method validates the JSON key path, loads the service-account file with service_account.Credentials.from_service_account_file(), and scopes it to https://www.googleapis.com/auth/analytics.readonly. It then passes those credentials to BetaAnalyticsDataClient, storing the resulting client in self.client.

Similarly, data_sources/modules/google_search_console.py implements GoogleSearchConsole.__init__ to load the same credential type but scopes it to https://www.googleapis.com/auth/webmasters.readonly. Instead of a dedicated client class, it builds a service object via googleapiclient.discovery.build('searchconsole', 'v1', credentials=...), exposing the underlying REST interface through self.service.

Google Analytics 4 Data Retrieval

Once authenticated, the GA4 module constructs typed protobuf requests to fetch reporting data.

Building the RunReportRequest

The get_top_pages method (and similar report methods) assembles a RunReportRequest object that specifies:

  • Property: The GA4 property ID prefixed with properties/
  • Date ranges: Relative strings such as "30daysAgo" to "today"
  • Dimensions: Fields like pagePath and pageTitle
  • Metrics: Quantities such as screenPageViews and sessions
  • Optional filters: Dimension filters to restrict results to specific path prefixes

Processing the Response

The method calls self.client.run_report(request) and iterates over the returned rows. It maps the protobuf dimension and metric values into Python dictionaries with human-friendly keys (e.g., pageviews, sessions, bounce_rate), returning a list of records ready for downstream analysis.

Google Search Console Data Retrieval

The GSC module follows a similar pattern but uses standard Python dictionaries to construct JSON payloads for the REST API.

Constructing the Query Body

Methods such as get_quick_wins build a request dictionary containing:

  • Start and end dates: ISO-formatted date strings or relative offsets
  • Dimensions: Typically ["query"] or ["page", "query"]
  • Row limits: Integer caps to control payload size
  • Dimension filters: Groups of filters to exclude branded terms or restrict to specific page paths

Executing and Parsing Results

The module invokes self.service.searchanalytics().query(siteUrl=self.site_url, body=request).execute(). The JSON response contains a rows list where each entry holds keys (dimension values) and metrics (clicks, impressions, ctr, position). The method sorts these rows by opportunity score or clicks and returns them as a list of standardized dictionaries.

Complete Implementation Examples

The following snippets demonstrate end-to-end usage, assuming you have exported the required environment variables.


# Fetch top-performing blog content from GA4

from data_sources.modules.google_analytics import GoogleAnalytics

ga = GoogleAnalytics()  # Authenticates automatically via env vars

top_blog_posts = ga.get_top_pages(
    days=30,
    limit=10,
    path_filter="/blog/"
)

for post in top_blog_posts:
    print(f"{post['title']}: {post['pageviews']:,} views")

# Identify quick-win keywords from Search Console

from data_sources.modules.google_search_console import GoogleSearchConsole

gsc = GoogleSearchConsole()  # Authenticates automatically via env vars

opportunities = gsc.get_quick_wins(
    days=30,
    position_min=11,
    position_max=20
)

for kw in opportunities[:5]:
    print(f"{kw['keyword']} (pos {kw['position']}) – "
          f"{kw['impressions']:,} impressions, score {kw['opportunity_score']}")

Summary

  • Service-account authentication keeps secrets out of source code by loading JSON keys via environment variables (GA4_CREDENTIALS_PATH, GSC_CREDENTIALS_PATH).
  • Google Analytics 4 uses the BetaAnalyticsDataClient with scoped credentials to send RunReportRequest protobufs and returns parsed dimension/metric dictionaries.
  • Google Search Console leverages googleapiclient.discovery.build to create a REST service object, executing searchanalytics().query() with Python dictionaries as request bodies.
  • Both modules encapsulate credential handling in their __init__ constructors, exposing simple high-level methods for downstream SEO analysis.

Frequently Asked Questions

What environment variables are required to run these modules?

You must set GA4_PROPERTY_ID and GA4_CREDENTIALS_PATH for the Analytics module, and GSC_SITE_URL and GSC_CREDENTIALS_PATH for the Search Console module. These variables point to your Google Cloud service-account JSON key files and target property or site identifiers.

Why do these modules use service accounts instead of OAuth 2.0 user flows?

Service accounts are designed for server-to-server interactions that do not require user intervention. Since seomachine runs as an automated data pipeline, service accounts allow the code to authenticate silently using a JSON key file without prompting for browser-based consent flows.

How does the GA4 client differ from the GSC client in terms of data retrieval?

The GA4 module uses the BetaAnalyticsDataClient, which accepts strongly-typed protobuf RunReportRequest objects and returns structured row data that is parsed into Python dictionaries. The GSC module uses the standard googleapiclient discovery interface, building plain Python dictionaries as JSON payloads and executing REST queries via searchanalytics().query().

Can I filter data before fetching it to reduce API quota usage?

Yes. Both modules support server-side filtering. In google_analytics.py, you can pass a path_filter argument to get_top_pages, which constructs a DimensionFilter inside the RunReportRequest. In google_search_console.py, methods like get_quick_wins accept filter parameters that become dimensionFilterGroups in the JSON request body, ensuring only relevant rows are returned.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →