How the Google integrations work: technical architecture

For technical evaluators: the polling, sync, identity and attribution pipelines behind Google Ads, Local Services and GA4, and what is stored where.

How the Google integrations work: technical architecture

This article is for the person who has read Google Ads & Local Services and Google Analytics 4 & Google Ads Conversions and now wants to know how. It describes the pipelines as they are built: what runs when, on which credentials, what each step reads from Google, what it writes on our side, and where the guarantees come from. Setup steps live in the how-to articles linked at the end; this one stays at the level of mechanism.

Three products share the machinery described here:

The pieces at a glance

flowchart LR
  subgraph G["Google"]
    ADS["Google Ads API<br/>campaigns · leads · clicks"]
    LSA["Local Services<br/>conversation relay"]
    GA4["GA4 property"]
    TAG["Your Google tag<br/>on your site"]
  end
  subgraph K["Karmaflow.ai"]
    POLL["Minute tick →<br/>per-account poller"]
    SYNC["Nightly lifecycle sync<br/>+ click resolution"]
    ENGINE["Shared SMS agent engine"]
    CRM[("CRM contact<br/>+ conversations")]
    FACTS[("Lifecycle facts<br/>per account · day · channel")]
    TRACK["Tracker · chat widget ·<br/>form endpoints"]
    BRIDGE["GA4 bridge +<br/>workflow steps"]
    READ["Lifecycle page ·<br/>agent tools · MCP"]
  end
  subgraph T["Your systems"]
    SCRM["Service CRM<br/>jobs + invoices"]
  end
  LSA -->|"GAQL read, every poll"| POLL
  POLL --> ENGINE
  ENGINE -->|"append reply"| LSA
  POLL --> CRM
  ADS -->|"GAQL reads, nightly"| SYNC
  SCRM -->|"outcomes"| SYNC
  SYNC --> FACTS
  SYNC -->|"ad context"| CRM
  TAG -->|"client id · session id · gclid"| TRACK
  TRACK --> CRM
  CRM --> BRIDGE
  BRIDGE -->|"Measurement Protocol"| GA4
  BRIDGE -.->|"offline conversion upload"| ADS
  FACTS --> READ
  CRM --> READ

Two cadences do different jobs. The poller runs against Local Services only, on a per-account interval, because Google offers no webhook for Local Services leads and the poll interval is therefore the reply latency. The nightly sync reads the whole ad account and materializes facts, so that every page view and every agent question afterwards costs zero Google operations. Nothing on our side reads Google when a user opens a page.

Access model

Two credentials, two axes

Developer token OAuth grant
Identifies Karmaflow.ai to the Google Ads API Karmaflow.ai to a Google account
Issued from Karmaflow.ai's Google Ads manager account Karmaflow.ai's Google Cloud project
Scales by operations per day, per token users, per project
Held where platform configuration, or a per-workspace secret if you bring your own encrypted on the integration record

The developer token carries Basic access: a daily budget of operations shared by every linked account on the platform, counted on a sliding 24-hour window rather than a calendar day. One read is one operation regardless of how many rows it returns; a request Google rejects still counts. That budget is why the fastest poll interval is a guarded setting, and why every read below is described with its operation cost.

The OAuth grant uses the Google Ads scope and is made by a Karmaflow.ai role account that has access to the manager account. The refresh token is stored encrypted at rest (AES-256-GCM); an access token is minted lazily and cached for about an hour. If Google invalidates the grant (revocation, six months unused, a rotated consent screen), the integration moves to a visible needs reconnect state and polling stops rather than retrying into the quota.

Reaching your account

Your Google Ads account is reached through the manager account, never by a grant on your own Google login:

  1. You enter your ten-digit account id. A manager-link invitation is sent to that account.
  2. You accept the invitation inside Google Ads.
  3. You press Verify. This is the only path to the linked state, and it is a real read: the account's name, time zone and currency are stamped on the integration, the campaign list is read to learn which channel types the account runs and whether it has a Local Services campaign, and a cold-start floor is stamped so that leads older than the connection are never auto-answered.

Every API call then carries the manager account in the login-customer-id header and your account id in the path. A workspace can link several accounts; they share one grant and are polled and synced independently.

What the credential may do

Two boundaries come from Google, not from us. Local Services campaigns cannot be created or edited through the API by anyone, and leads in healthcare categories are never returned by the API at all, so a healthcare-category advertiser gets the messaging channel's no data state rather than a degraded one.

Transport and versioning

Calls go over HTTPS to Google's REST endpoints; there is no client library in the path. The major API version is pinned and each major lives roughly a year, so the maintenance cadence is one deliberate upgrade per year rather than a quarterly scramble.

Pipeline 1 — Local Services message leads (near-real time)

sequenceDiagram
  participant S as Cloud Scheduler
  participant P as Poller
  participant G as Google Ads API
  participant L as Lead record + CRM contact
  participant A as SMS agent engine
  participant C as Consumer
  S->>P: tick, every minute
  P->>P: is this account due? (its own interval)
  P->>G: one GAQL read: conversations since cursor − 5 min
  G-->>P: consumer rows, advertiser rows, lead fields
  P->>P: classify each row
  P->>L: upsert lead, resolve or create contact
  P->>A: hand over the consumer message (thread = the lead)
  A-->>P: reply, or a draft held for approval
  P->>G: append reply to the lead conversation
  G-->>C: delivered on the consumer's channel
  Note over P,G: next poll: our reply returns as an advertiser row and is recognised as ours

Cadence

A scheduler hits the tick endpoint every minute, authenticated by a shared secret. The tick sweeps workspaces in batches and asks each linked account whether it is due. An account polls on its own interval: default 120 seconds, floor 60, ceiling 3,600. Accounts that are switched off, not yet linked, in a needs reconnect state, without a Local Services campaign, or inside a backoff window are skipped with a named reason that the tick reports.

The read

One query per poll, one operation. It selects from the lead-conversation resource joined to its lead: conversation id, the lead's resource name, event time, participant type, channel, message text and attachment URLs, plus the lead's status, type, category, locale, charge state, credit state and contact details. It is filtered to message-type leads and to events after the cursor.

The cursor is kept in the ad account's local time, because that is how Google stamps conversation times; comparing it against a UTC clock would skip or re-fetch by the account's offset. Each poll steps the cursor back five minutes so nothing is lost at the boundary, which guarantees re-delivery and therefore makes idempotency mandatory (below).

The query deliberately does not filter to consumer rows. Advertiser-side rows are fetched too, because a row on the business side that we did not write is a person replying in Google's own dashboard, and the takeover feature needs to see it.

Idempotency, in three layers

  1. A per-account set of recently seen conversation ids, checked before anything else.
  2. A unique index on the conversation id stored on the message row; a duplicate insert is recognised and produces no reply.
  3. The cursor advances only when a row succeeds or is skipped on purpose. A transient failure halts the batch without moving the cursor, so the message is retried on the next poll.

Row classification

Row What happens
Lead in a terminal status (expired, disabled, declined by either side, wiped out) Lead record updated, no reply ever attempted. Wiped out means Google erased the consumer's details; the lead is stamped with an erasure timestamp and no contact is created from it.
Advertiser row with a conversation id we recorded when we replied, or on the ADS_API channel Ours. Ingested, never treated as inbound. An exact-text check against our own sent messages catches an echo the first two tests missed.
Any other advertiser row A person replied in Google's dashboard. The account's dashboard-reply policy is applied to that lead unless a person already chose a per-lead setting. Counts are kept on the integration so the settings page can show what was seen.
Consumer row on a channel the agent does not answer (email, WhatsApp by default) Recorded and visible in the thread; no reply.
Consumer row on an answered channel Delivered to the agent.

Lead record and CRM contact

Every row upserts a lead record keyed on Google's lead id. It carries Google's fields (status, type, category, charge and credit state, creation time, channels seen), the consumer's name and phone as Google supplied them, and our own state: message counts, first and last reply timestamps and who replied, the per-lead AI handling, opt-out, cap and takeover flags, and the conversation ids of our own replies.

A CRM contact is resolved or created once per lead:

Gates before the agent runs

In order: STOP / START keywords (per lead, and optionally recorded as a platform opt-out event); the per-lead AI handling set by a person (keep replying, draft for me, I'll take it from here); the fair-use cap of 30 AI messages per lead, after which the draft queues for a person and admins are notified. A lead that predates the cold-start floor is never reached by the poller at all; it is imported only by the explicit Import past leads action, which replies to nothing.

The shared conversation engine

The message is handed to the same engine that answers ordinary texts on the SMS channel, with the transport swapped:

Delivery and provenance

Delivery is one call: append the reply to the lead's conversation. The response is inspected per item, because Google reports a rejected append inside a successful HTTP response; a failed append is never recorded as sent. On failure, an optional fallback sends the same text as an SMS from a number you configure, only when the consumer's number is unique to one lead and the contact is not opted out of SMS. The fallback is off by default so that a failing Google channel is visible rather than papered over.

After delivery: Google's conversation id for our reply is recorded on the lead (that is how the next poll recognises it as ours); first-reply and last-reply timestamps are stamped with who replied (agent, or a named user); a sent-message record is written with provider Google Local Services and links back to the agent, thread, contact and any workflow that produced it. The first delivered reply on a lead claims one Local Services thread billing unit; the claim is atomic, and a failed send rolls it back.

Backoff

Failure Backoff
Temporary rate limit 30 seconds, not counted as a failure
Daily quota exhausted 1 hour, so the sliding window recovers instead of burning more rejected requests
Any other error Exponential from 1 minute, capped at 30 minutes
Three delivery failures on one message The lead is flagged for a person and admins are notified; the queue moves on
Invalid grant Needs reconnect; polling stops until a person reconnects

Google first, for as long as the lead is live

Any SMS the platform sends to a consumer whose lead is live (a reply is possible, the consumer has not opted out, and they were heard from within 30 days) is routed through Google's conversation instead of the carrier, whether it comes from a workflow, a campaign or an agent tool. The lead is chosen by the contact, never by a phone number shared across leads. A separate guard at the single SMS send function refuses any carrier text to a number seen on two or more leads, because a tracking number does not reach the consumer and can be reassigned after the lead closes.

Pipeline 2 — Search, Performance Max and Display leads

Google keeps no conversation for these campaign types, so what they produce reaches Karmaflow.ai through your own numbers and forms, and the platform works out that Google sent it without asking Google.

Source How it arrives How it is attributed
Message asset on a Search or Performance Max ad An ordinary text from the consumer's real phone to your SMS agent's number The starter message you copied into settings, or a phone number flagged as a Google Ads destination
Call asset, or the Local Services call destination A call to a Karmaflow.ai number answered by a voice agent The landing number, flagged as a Google Ads destination. Google passes no click id on a forwarded call.
Lead form on Search, Performance Max, Video or Display Polled from the lead-form submission resource, one operation per poll, every 5 minutes by default The submission carries campaign, ad group, ad and click id

Lead forms get a dedicated poller on the same credential: submission times carry an explicit UTC offset, so the cursor is a true instant with a ten-minute overlap and de-duplication on submission id. Each submission becomes a form-lead record, a contact (found by email first, then phone — both were typed by the consumer, so identity is trustworthy), a follow-up task and a notification. The click id is queued for resolution (Pipeline 4). Accounts that run only a Local Services campaign are skipped, since lead forms cannot attach to one.

Pipeline 3 — Nightly lifecycle sync

The sync runs once a night per linked account with the lifecycle sync enabled, or on demand from Sync now. It re-reads a trailing window of 7 account-local days (configurable, 1–90) every time, because Google restates the past: a lead's credit state moves from pending to credited days later, which changes what a finished day cost.

Capability first

Before any query, the sync asks Google's own field-metadata service which fields are selectable on this API version and assembles its selects from the answer. The result is cached per API version, not per account. A short list of fields Google has rejected live outranks the metadata and the fallback alike, so a metadata outage can never replay a known failure. What cannot be selected becomes a capability the page reports (this API version does not expose cost) rather than an error a user discovers.

The reads

Step Reads from Google Cost Produces
Spend Campaign metrics by date: cost, impressions, clicks, conversions, conversion value, view-through, impression share and lost share, account currency 1 op One fact row per (account, day, channel type)
Engagements Every Local Services lead in the window: kind, status, charged, credited 1 op Lead counts by kind and status, charged and credited counts
Calls Lead conversations with phone-call duration and recording URL 1 op Call count, talk time, recorded count
Search terms Top search terms for Search campaigns; Google's search-term themes 1–2 ops Term tables on the account and on the last day's Search row
Verification License and insurance verification artifacts with expiry 1 op The lapsed / expiring warnings
Budgets and changes Campaign budgets; the account's change history (30 days) 2 ops The budget card and the change list

Steps fail independently: a rejected read records its reason on the integration and the rest of the sync continues. Local Services reads are skipped on accounts that have no Local Services campaign, because an empty read still costs an operation. A typical account costs 5–8 operations per night.

Our side of the join

Two more steps read nothing from Google. The lead records are bucketed by the lead's own day to compute the answered rate, who answered, and the median first response measured from Google's timestamp. Then outcomes are fetched from the connected service CRM through an outcomes port (one adapter per connector) and matched to leads by the attribution ladder:

Rung Meaning Deterministic
Stamped source The job carries the lead's own source identifier, written when the job was created Yes
Confirmed by a person Someone linked them by hand; never overwritten by a later pass Yes
Email match Exact match on an email the consumer volunteered in conversation No
Phone match Phone plus a 30-day window; numbers seen on more than one lead are excluded No

Matching is re-run over open outcome rows every night and promotes a match in place (a lead that matched by phone yesterday and gained an email today becomes an email match), never appending a second row, so revenue is never double-counted. Coverage is reported per rung, and unattributed revenue is counted beside it rather than hidden.

The fact rows

One row per (account, account-local day, channel type), upserted idempotently. Three rules are enforced by the schema:

Each row also carries a compute version. When a calculation changes in a release, rows written by the older version are flagged on the page until a sync recomputes them.

Reading the facts

The lifecycle page, the agent tools and the MCP tools read only these rows and the snapshots on the integration. Opening the page, asking an agent, or calling MCP costs no Google operations; the figures are as fresh as the last sync, and the page says when that was.

Pipeline 4 — Click → visitor → contact → conversation

sequenceDiagram
  participant V as Visitor's browser
  participant T as Your Google tag
  participant K as Tracker · chat widget · form
  participant C as Contact record
  participant R as Nightly click resolution
  participant G as Google Ads API
  V->>T: lands from an ad (gclid in the URL)
  K->>T: gtag('get') → client id, session id
  V->>K: chats, submits a form, or identifies
  K->>C: client id, session id, gclid/gbraid/wbraid, UTM, landing path
  R->>G: click_view for the click's day (one op per account per day)
  G-->>R: campaign, ad group, keyword, ad
  R->>G: ad_group_ad for the ads seen (one op)
  R->>C: ad context on the contact and on its conversations

Capture

A visitor becomes attributable at an identification moment while a Karmaflow.ai surface and your Google tag are both on the page: opening the chat widget, submitting a form that posts to a Karmaflow.ai form endpoint, or the tracker's identify call. The surface reads the GA client id and session id from your tag using its documented API (with a cookie fallback), and the click ids and UTM values from the landing URL; the tracker keeps click ids for 90 days so a chat days after the click still carries them, and injects them as hidden fields into forms.

On the contact, the GA block is most-recent-wins (the newest browser is where the next click will come from) while click ids and UTM keep the first capture (the acquisition touch). The chat session keeps its own copy, and the contact adopts it when the session is linked to a person. Emails, phones, names and message text never travel in these fields.

Resolution

Google exposes which campaign, ad group, keyword and ad a click id came from, with two constraints: a query must name a single day, and the data reaches back 90 days. So each captured click id becomes a pending row, and every night each account that runs click-landing campaigns (Search, Performance Max, Display, Video, Demand Gen, Shopping — never Local Services, whose leads carry no click id) is asked about the few days its pending rows point at, at most seven days per account per night. Because a day's answer lists every click that day, a pending id absent from it is known not to have clicked that day, which makes exhaustion provable: rows run out of days and are retired with a stated reason. One further read fetches the ads' headlines, descriptions and final URLs, cached per ad.

The resolved ad is written onto the contact and copied onto every conversation the contact has (chat, SMS, call, email), at resolution time and again when a new conversation is analyzed. That is what lets every insight be sliced by campaign, ad group or keyword without a new query engine.

What the join enables

Each conversation is analyzed minutes after it ends into a structured envelope: intents, outcome type, objections and whether they were resolved, drop-off reason, the questions the agent could not answer verbatim, sentiment. With the ad context stamped on it, the lifecycle page can show what the people a keyword brought actually said, and set the terms your ads were clicked for beside the topics people raised once they arrived. The comparison is deterministic (shared significant tokens after stop-word removal) and auditable; where AI-written findings are added on top, a second model on a different provider reviews them and any sentence whose number is not in the computed figures is dropped before publication. Nothing is ever applied to your Google Ads account from this report.

Pipeline 5 — Closing the loop to Google

GA4 via the Measurement Protocol

Sending to GA4 needs no OAuth and nothing installed beyond the tag you already have: you paste a Measurement ID and an API secret (stored encrypted), and Karmaflow.ai's servers send events through Google's Measurement Protocol, to the global or the EU endpoint as you choose.

Offline conversions into Google Ads

The direct path reports a conversion to your Google Ads account by click id, or by hashed email and phone (Enhanced Conversions for Leads) when there is no click id. It is a write, so it is gated four times, in order: a platform-wide switch, your workspace's approval for Google data sharing, the account's own switch, and a named conversion action of the accepting type. Uploads happen only from a workflow step you build; nothing uploads automatically.

The click id is looked for in order: the run itself, the contact, the contact's most recent chat session with a click id, a Google lead-form lead. Each upload carries an order id (by default the trigger and event id, or a field you map such as the deal id), so a re-run cannot double-count. Google answers per row, and each attempt is logged with its result — accepted, already recorded, rejected with Google's reason, or not sent with the gate that stopped it — for 90 days, with the click id stored only as a hash.

Reading GA4 back

Where it is enabled for your workspace, one administrator's Google grant with the read-only Analytics scope lets Karmaflow.ai read aggregate reports from the property you send to: landing pages by channel, the paid-search queries GA4 credits, event counts by name and day, and a reconciliation of what we sent against what GA4 recorded. Reads are nightly snapshots with a probe of the property's own metadata first; nothing reads GA4 on a page view, and Google's refusals are stored beside each report.

First-party web analytics, briefly

The same tracker that captures identity is a full first-party analytics collector. Events are posted in batches to a collection endpoint (CORS, origin allow-list, bot filtering, IP hashed after geo lookup), streamed into a BigQuery events table partitioned by day and clustered by workspace, site and event name, and rolled up into per-workspace collections for dashboards. A tag manager compiles tags, triggers and variables into a versioned container served beside the tracker; one tag type fires events through your own Google tag on the page. The full description is in Web Analytics.

What is stored, where, and for how long

Every record below lives in your workspace's own database unless stated.

Record Holds Retention
Google Ads integration Account ids, link state, encrypted grant, time zone, currency, cursor, behaviour settings, quota counters, nightly snapshots (search terms, verification, budgets, change history) Life of the connection
Lead record Google's lead fields, consumer name and phone as supplied, typed identity, reply state, AI handling, erasure and billing stamps Kept; erasure stamp on wipe-out
Message thread and messages The conversation, attachments, who wrote each turn, approval state Your transcript retention policy
Sent-message provenance Every reply with provider, agent or user, thread, contact and workflow links Kept
Lifecycle facts and lead outcomes Per-day facts with tier and compute version; attributed outcomes with rung and evidence Kept; re-synced over the trailing window
Ad clicks and creatives Click id → campaign, ad group, keyword, ad; cached ad text 455 days
Form leads Lead-form submissions with campaign, ad group, ad and click id Kept
Contact analytics identity GA client id, session id, click ids, UTM, landing path, resolved ad Life of the contact
GA4 integration Measurement ID, encrypted API secret, region, toggles, event map Life of the connection
GA4 delivery log One row per send or skip, with reason; client id as a hash only 30 days
Conversion uploads One row per attempt with result and Google's reason; click id hashed 90 days
GA4 report snapshots Aggregate rows per report and property Replaced nightly

Tenancy, security and privacy

Operational characteristics

Boundaries, stated plainly

Learn more

Sign in to Karmaflow.ai