How to Authenticate and Fetch Data from Google Analytics 4 and Search Console in Python
Both modules authenticate via service-account JSON keys stored in environment variables, then fetch data using the BetaAnalyticsDataClient for GA4 and the googleapiclient discovery service for Search Console.
The seomachine repository provides Python wrappers that simplify extracting SEO metrics from Google’s APIs. Understanding how google_analytics.py and google_search_console.py handle authentication and data retrieval lets you integrate these data sources securely into your own pipelines without hardcoding secrets.
Service-Account Authentication Architecture
Both modules follow an identical service-account pattern: they load a Google Cloud JSON key from disk, create scoped credentials, and initialize a client object that persists for the lifetime of the class instance.
Environment Variable Configuration
Secrets are externalized to environment variables to keep the codebase free of credentials. The modules read the following keys during initialization:
google_analytics.py:GA4_PROPERTY_IDandGA4_CREDENTIALS_PATHgoogle_search_console.py:GSC_SITE_URLandGSC_CREDENTIALS_PATH
Credential Loading and Client Initialization
In data_sources/modules/google_analytics.py, the GoogleAnalytics.__init__ method validates the JSON key path, loads the service-account file with service_account.Credentials.from_service_account_file(), and scopes it to https://www.googleapis.com/auth/analytics.readonly. It then passes those credentials to BetaAnalyticsDataClient, storing the resulting client in self.client.
Similarly, data_sources/modules/google_search_console.py implements GoogleSearchConsole.__init__ to load the same credential type but scopes it to https://www.googleapis.com/auth/webmasters.readonly. Instead of a dedicated client class, it builds a service object via googleapiclient.discovery.build('searchconsole', 'v1', credentials=...), exposing the underlying REST interface through self.service.
Google Analytics 4 Data Retrieval
Once authenticated, the GA4 module constructs typed protobuf requests to fetch reporting data.
Building the RunReportRequest
The get_top_pages method (and similar report methods) assembles a RunReportRequest object that specifies:
- Property: The GA4 property ID prefixed with
properties/ - Date ranges: Relative strings such as
"30daysAgo"to"today" - Dimensions: Fields like
pagePathandpageTitle - Metrics: Quantities such as
screenPageViewsandsessions - Optional filters: Dimension filters to restrict results to specific path prefixes
Processing the Response
The method calls self.client.run_report(request) and iterates over the returned rows. It maps the protobuf dimension and metric values into Python dictionaries with human-friendly keys (e.g., pageviews, sessions, bounce_rate), returning a list of records ready for downstream analysis.
Google Search Console Data Retrieval
The GSC module follows a similar pattern but uses standard Python dictionaries to construct JSON payloads for the REST API.
Constructing the Query Body
Methods such as get_quick_wins build a request dictionary containing:
- Start and end dates: ISO-formatted date strings or relative offsets
- Dimensions: Typically
["query"]or["page", "query"] - Row limits: Integer caps to control payload size
- Dimension filters: Groups of filters to exclude branded terms or restrict to specific page paths
Executing and Parsing Results
The module invokes self.service.searchanalytics().query(siteUrl=self.site_url, body=request).execute(). The JSON response contains a rows list where each entry holds keys (dimension values) and metrics (clicks, impressions, ctr, position). The method sorts these rows by opportunity score or clicks and returns them as a list of standardized dictionaries.
Complete Implementation Examples
The following snippets demonstrate end-to-end usage, assuming you have exported the required environment variables.
# Fetch top-performing blog content from GA4
from data_sources.modules.google_analytics import GoogleAnalytics
ga = GoogleAnalytics() # Authenticates automatically via env vars
top_blog_posts = ga.get_top_pages(
days=30,
limit=10,
path_filter="/blog/"
)
for post in top_blog_posts:
print(f"{post['title']}: {post['pageviews']:,} views")
# Identify quick-win keywords from Search Console
from data_sources.modules.google_search_console import GoogleSearchConsole
gsc = GoogleSearchConsole() # Authenticates automatically via env vars
opportunities = gsc.get_quick_wins(
days=30,
position_min=11,
position_max=20
)
for kw in opportunities[:5]:
print(f"{kw['keyword']} (pos {kw['position']}) – "
f"{kw['impressions']:,} impressions, score {kw['opportunity_score']}")
Summary
- Service-account authentication keeps secrets out of source code by loading JSON keys via environment variables (
GA4_CREDENTIALS_PATH,GSC_CREDENTIALS_PATH). - Google Analytics 4 uses the
BetaAnalyticsDataClientwith scoped credentials to sendRunReportRequestprotobufs and returns parsed dimension/metric dictionaries. - Google Search Console leverages
googleapiclient.discovery.buildto create a REST service object, executingsearchanalytics().query()with Python dictionaries as request bodies. - Both modules encapsulate credential handling in their
__init__constructors, exposing simple high-level methods for downstream SEO analysis.
Frequently Asked Questions
What environment variables are required to run these modules?
You must set GA4_PROPERTY_ID and GA4_CREDENTIALS_PATH for the Analytics module, and GSC_SITE_URL and GSC_CREDENTIALS_PATH for the Search Console module. These variables point to your Google Cloud service-account JSON key files and target property or site identifiers.
Why do these modules use service accounts instead of OAuth 2.0 user flows?
Service accounts are designed for server-to-server interactions that do not require user intervention. Since seomachine runs as an automated data pipeline, service accounts allow the code to authenticate silently using a JSON key file without prompting for browser-based consent flows.
How does the GA4 client differ from the GSC client in terms of data retrieval?
The GA4 module uses the BetaAnalyticsDataClient, which accepts strongly-typed protobuf RunReportRequest objects and returns structured row data that is parsed into Python dictionaries. The GSC module uses the standard googleapiclient discovery interface, building plain Python dictionaries as JSON payloads and executing REST queries via searchanalytics().query().
Can I filter data before fetching it to reduce API quota usage?
Yes. Both modules support server-side filtering. In google_analytics.py, you can pass a path_filter argument to get_top_pages, which constructs a DimensionFilter inside the RunReportRequest. In google_search_console.py, methods like get_quick_wins accept filter parameters that become dimensionFilterGroups in the JSON request body, ensuring only relevant rows are returned.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →