For technical evaluators: the polling, sync, identity and attribution pipelines behind Google Ads, Local Services and GA4, and what is stored where.
This article is for the person who has read Google Ads & Local Services and Google Analytics 4 & Google Ads Conversions and now wants to know how. It describes the pipelines as they are built: what runs when, on which credentials, what each step reads from Google, what it writes on our side, and where the guarantees come from. Setup steps live in the how-to articles linked at the end; this one stays at the level of mechanism.
Three products share the machinery described here:
flowchart LR
subgraph G["Google"]
ADS["Google Ads API<br/>campaigns · leads · clicks"]
LSA["Local Services<br/>conversation relay"]
GA4["GA4 property"]
TAG["Your Google tag<br/>on your site"]
end
subgraph K["Karmaflow.ai"]
POLL["Minute tick →<br/>per-account poller"]
SYNC["Nightly lifecycle sync<br/>+ click resolution"]
ENGINE["Shared SMS agent engine"]
CRM[("CRM contact<br/>+ conversations")]
FACTS[("Lifecycle facts<br/>per account · day · channel")]
TRACK["Tracker · chat widget ·<br/>form endpoints"]
BRIDGE["GA4 bridge +<br/>workflow steps"]
READ["Lifecycle page ·<br/>agent tools · MCP"]
end
subgraph T["Your systems"]
SCRM["Service CRM<br/>jobs + invoices"]
end
LSA -->|"GAQL read, every poll"| POLL
POLL --> ENGINE
ENGINE -->|"append reply"| LSA
POLL --> CRM
ADS -->|"GAQL reads, nightly"| SYNC
SCRM -->|"outcomes"| SYNC
SYNC --> FACTS
SYNC -->|"ad context"| CRM
TAG -->|"client id · session id · gclid"| TRACK
TRACK --> CRM
CRM --> BRIDGE
BRIDGE -->|"Measurement Protocol"| GA4
BRIDGE -.->|"offline conversion upload"| ADS
FACTS --> READ
CRM --> READ
Two cadences do different jobs. The poller runs against Local Services only, on a per-account interval, because Google offers no webhook for Local Services leads and the poll interval is therefore the reply latency. The nightly sync reads the whole ad account and materializes facts, so that every page view and every agent question afterwards costs zero Google operations. Nothing on our side reads Google when a user opens a page.
| Developer token | OAuth grant | |
|---|---|---|
| Identifies | Karmaflow.ai to the Google Ads API | Karmaflow.ai to a Google account |
| Issued from | Karmaflow.ai's Google Ads manager account | Karmaflow.ai's Google Cloud project |
| Scales by | operations per day, per token | users, per project |
| Held where | platform configuration, or a per-workspace secret if you bring your own | encrypted on the integration record |
The developer token carries Basic access: a daily budget of operations shared by every linked account on the platform, counted on a sliding 24-hour window rather than a calendar day. One read is one operation regardless of how many rows it returns; a request Google rejects still counts. That budget is why the fastest poll interval is a guarded setting, and why every read below is described with its operation cost.
The OAuth grant uses the Google Ads scope and is made by a Karmaflow.ai role account that has access to the manager account. The refresh token is stored encrypted at rest (AES-256-GCM); an access token is minted lazily and cached for about an hour. If Google invalidates the grant (revocation, six months unused, a rotated consent screen), the integration moves to a visible needs reconnect state and polling stops rather than retrying into the quota.
Your Google Ads account is reached through the manager account, never by a grant on your own Google login:
Every API call then carries the manager account in the login-customer-id header and your account id in the path. A workspace can link several accounts; they share one grant and are polled and synced independently.
Two boundaries come from Google, not from us. Local Services campaigns cannot be created or edited through the API by anyone, and leads in healthcare categories are never returned by the API at all, so a healthcare-category advertiser gets the messaging channel's no data state rather than a degraded one.
Calls go over HTTPS to Google's REST endpoints; there is no client library in the path. The major API version is pinned and each major lives roughly a year, so the maintenance cadence is one deliberate upgrade per year rather than a quarterly scramble.
sequenceDiagram
participant S as Cloud Scheduler
participant P as Poller
participant G as Google Ads API
participant L as Lead record + CRM contact
participant A as SMS agent engine
participant C as Consumer
S->>P: tick, every minute
P->>P: is this account due? (its own interval)
P->>G: one GAQL read: conversations since cursor − 5 min
G-->>P: consumer rows, advertiser rows, lead fields
P->>P: classify each row
P->>L: upsert lead, resolve or create contact
P->>A: hand over the consumer message (thread = the lead)
A-->>P: reply, or a draft held for approval
P->>G: append reply to the lead conversation
G-->>C: delivered on the consumer's channel
Note over P,G: next poll: our reply returns as an advertiser row and is recognised as ours
A scheduler hits the tick endpoint every minute, authenticated by a shared secret. The tick sweeps workspaces in batches and asks each linked account whether it is due. An account polls on its own interval: default 120 seconds, floor 60, ceiling 3,600. Accounts that are switched off, not yet linked, in a needs reconnect state, without a Local Services campaign, or inside a backoff window are skipped with a named reason that the tick reports.
One query per poll, one operation. It selects from the lead-conversation resource joined to its lead: conversation id, the lead's resource name, event time, participant type, channel, message text and attachment URLs, plus the lead's status, type, category, locale, charge state, credit state and contact details. It is filtered to message-type leads and to events after the cursor.
The cursor is kept in the ad account's local time, because that is how Google stamps conversation times; comparing it against a UTC clock would skip or re-fetch by the account's offset. Each poll steps the cursor back five minutes so nothing is lost at the boundary, which guarantees re-delivery and therefore makes idempotency mandatory (below).
The query deliberately does not filter to consumer rows. Advertiser-side rows are fetched too, because a row on the business side that we did not write is a person replying in Google's own dashboard, and the takeover feature needs to see it.
| Row | What happens |
|---|---|
| Lead in a terminal status (expired, disabled, declined by either side, wiped out) | Lead record updated, no reply ever attempted. Wiped out means Google erased the consumer's details; the lead is stamped with an erasure timestamp and no contact is created from it. |
Advertiser row with a conversation id we recorded when we replied, or on the ADS_API channel |
Ours. Ingested, never treated as inbound. An exact-text check against our own sent messages catches an echo the first two tests missed. |
| Any other advertiser row | A person replied in Google's dashboard. The account's dashboard-reply policy is applied to that lead unless a person already chose a per-lead setting. Counts are kept on the integration so the settings page can show what was seen. |
| Consumer row on a channel the agent does not answer (email, WhatsApp by default) | Recorded and visible in the thread; no reply. |
| Consumer row on an answered channel | Delivered to the agent. |
Every row upserts a lead record keyed on Google's lead id. It carries Google's fields (status, type, category, charge and credit state, creation time, channels seen), the consumer's name and phone as Google supplied them, and our own state: message counts, first and last reply timestamps and who replied, the per-lead AI handling, opt-out, cap and takeover flags, and the conversation ids of our own replies.
A CRM contact is resolved or created once per lead:
In order: STOP / START keywords (per lead, and optionally recorded as a platform opt-out event); the per-lead AI handling set by a person (keep replying, draft for me, I'll take it from here); the fair-use cap of 30 AI messages per lead, after which the draft queues for a person and admins are notified. A lead that predates the cold-start floor is never reached by the poller at all; it is imported only by the explicit Import past leads action, which replies to nothing.
The message is handed to the same engine that answers ordinary texts on the SMS channel, with the transport swapped:
Delivery is one call: append the reply to the lead's conversation. The response is inspected per item, because Google reports a rejected append inside a successful HTTP response; a failed append is never recorded as sent. On failure, an optional fallback sends the same text as an SMS from a number you configure, only when the consumer's number is unique to one lead and the contact is not opted out of SMS. The fallback is off by default so that a failing Google channel is visible rather than papered over.
After delivery: Google's conversation id for our reply is recorded on the lead (that is how the next poll recognises it as ours); first-reply and last-reply timestamps are stamped with who replied (agent, or a named user); a sent-message record is written with provider Google Local Services and links back to the agent, thread, contact and any workflow that produced it. The first delivered reply on a lead claims one Local Services thread billing unit; the claim is atomic, and a failed send rolls it back.
| Failure | Backoff |
|---|---|
| Temporary rate limit | 30 seconds, not counted as a failure |
| Daily quota exhausted | 1 hour, so the sliding window recovers instead of burning more rejected requests |
| Any other error | Exponential from 1 minute, capped at 30 minutes |
| Three delivery failures on one message | The lead is flagged for a person and admins are notified; the queue moves on |
| Invalid grant | Needs reconnect; polling stops until a person reconnects |
Any SMS the platform sends to a consumer whose lead is live (a reply is possible, the consumer has not opted out, and they were heard from within 30 days) is routed through Google's conversation instead of the carrier, whether it comes from a workflow, a campaign or an agent tool. The lead is chosen by the contact, never by a phone number shared across leads. A separate guard at the single SMS send function refuses any carrier text to a number seen on two or more leads, because a tracking number does not reach the consumer and can be reassigned after the lead closes.
Google keeps no conversation for these campaign types, so what they produce reaches Karmaflow.ai through your own numbers and forms, and the platform works out that Google sent it without asking Google.
| Source | How it arrives | How it is attributed |
|---|---|---|
| Message asset on a Search or Performance Max ad | An ordinary text from the consumer's real phone to your SMS agent's number | The starter message you copied into settings, or a phone number flagged as a Google Ads destination |
| Call asset, or the Local Services call destination | A call to a Karmaflow.ai number answered by a voice agent | The landing number, flagged as a Google Ads destination. Google passes no click id on a forwarded call. |
| Lead form on Search, Performance Max, Video or Display | Polled from the lead-form submission resource, one operation per poll, every 5 minutes by default | The submission carries campaign, ad group, ad and click id |
Lead forms get a dedicated poller on the same credential: submission times carry an explicit UTC offset, so the cursor is a true instant with a ten-minute overlap and de-duplication on submission id. Each submission becomes a form-lead record, a contact (found by email first, then phone — both were typed by the consumer, so identity is trustworthy), a follow-up task and a notification. The click id is queued for resolution (Pipeline 4). Accounts that run only a Local Services campaign are skipped, since lead forms cannot attach to one.
The sync runs once a night per linked account with the lifecycle sync enabled, or on demand from Sync now. It re-reads a trailing window of 7 account-local days (configurable, 1–90) every time, because Google restates the past: a lead's credit state moves from pending to credited days later, which changes what a finished day cost.
Before any query, the sync asks Google's own field-metadata service which fields are selectable on this API version and assembles its selects from the answer. The result is cached per API version, not per account. A short list of fields Google has rejected live outranks the metadata and the fallback alike, so a metadata outage can never replay a known failure. What cannot be selected becomes a capability the page reports (this API version does not expose cost) rather than an error a user discovers.
| Step | Reads from Google | Cost | Produces |
|---|---|---|---|
| Spend | Campaign metrics by date: cost, impressions, clicks, conversions, conversion value, view-through, impression share and lost share, account currency | 1 op | One fact row per (account, day, channel type) |
| Engagements | Every Local Services lead in the window: kind, status, charged, credited | 1 op | Lead counts by kind and status, charged and credited counts |
| Calls | Lead conversations with phone-call duration and recording URL | 1 op | Call count, talk time, recorded count |
| Search terms | Top search terms for Search campaigns; Google's search-term themes | 1–2 ops | Term tables on the account and on the last day's Search row |
| Verification | License and insurance verification artifacts with expiry | 1 op | The lapsed / expiring warnings |
| Budgets and changes | Campaign budgets; the account's change history (30 days) | 2 ops | The budget card and the change list |
Steps fail independently: a rejected read records its reason on the integration and the rest of the sync continues. Local Services reads are skipped on accounts that have no Local Services campaign, because an empty read still costs an operation. A typical account costs 5–8 operations per night.
Two more steps read nothing from Google. The lead records are bucketed by the lead's own day to compute the answered rate, who answered, and the median first response measured from Google's timestamp. Then outcomes are fetched from the connected service CRM through an outcomes port (one adapter per connector) and matched to leads by the attribution ladder:
| Rung | Meaning | Deterministic |
|---|---|---|
| Stamped source | The job carries the lead's own source identifier, written when the job was created | Yes |
| Confirmed by a person | Someone linked them by hand; never overwritten by a later pass | Yes |
| Email match | Exact match on an email the consumer volunteered in conversation | No |
| Phone match | Phone plus a 30-day window; numbers seen on more than one lead are excluded | No |
Matching is re-run over open outcome rows every night and promotes a match in place (a lead that matched by phone yesterday and gained an email today becomes an email match), never appending a second row, so revenue is never double-counted. Coverage is reported per rung, and unattributed revenue is counted beside it rather than hidden.
One row per (account, account-local day, channel type), upserted idempotently. Three rules are enforced by the schema:
Each row also carries a compute version. When a calculation changes in a release, rows written by the older version are flagged on the page until a sync recomputes them.
The lifecycle page, the agent tools and the MCP tools read only these rows and the snapshots on the integration. Opening the page, asking an agent, or calling MCP costs no Google operations; the figures are as fresh as the last sync, and the page says when that was.
sequenceDiagram
participant V as Visitor's browser
participant T as Your Google tag
participant K as Tracker · chat widget · form
participant C as Contact record
participant R as Nightly click resolution
participant G as Google Ads API
V->>T: lands from an ad (gclid in the URL)
K->>T: gtag('get') → client id, session id
V->>K: chats, submits a form, or identifies
K->>C: client id, session id, gclid/gbraid/wbraid, UTM, landing path
R->>G: click_view for the click's day (one op per account per day)
G-->>R: campaign, ad group, keyword, ad
R->>G: ad_group_ad for the ads seen (one op)
R->>C: ad context on the contact and on its conversations
A visitor becomes attributable at an identification moment while a Karmaflow.ai surface and your Google tag are both on the page: opening the chat widget, submitting a form that posts to a Karmaflow.ai form endpoint, or the tracker's identify call. The surface reads the GA client id and session id from your tag using its documented API (with a cookie fallback), and the click ids and UTM values from the landing URL; the tracker keeps click ids for 90 days so a chat days after the click still carries them, and injects them as hidden fields into forms.
On the contact, the GA block is most-recent-wins (the newest browser is where the next click will come from) while click ids and UTM keep the first capture (the acquisition touch). The chat session keeps its own copy, and the contact adopts it when the session is linked to a person. Emails, phones, names and message text never travel in these fields.
Google exposes which campaign, ad group, keyword and ad a click id came from, with two constraints: a query must name a single day, and the data reaches back 90 days. So each captured click id becomes a pending row, and every night each account that runs click-landing campaigns (Search, Performance Max, Display, Video, Demand Gen, Shopping — never Local Services, whose leads carry no click id) is asked about the few days its pending rows point at, at most seven days per account per night. Because a day's answer lists every click that day, a pending id absent from it is known not to have clicked that day, which makes exhaustion provable: rows run out of days and are retired with a stated reason. One further read fetches the ads' headlines, descriptions and final URLs, cached per ad.
The resolved ad is written onto the contact and copied onto every conversation the contact has (chat, SMS, call, email), at resolution time and again when a new conversation is analyzed. That is what lets every insight be sliced by campaign, ad group or keyword without a new query engine.
Each conversation is analyzed minutes after it ends into a structured envelope: intents, outcome type, objections and whether they were resolved, drop-off reason, the questions the agent could not answer verbatim, sentiment. With the ad context stamped on it, the lifecycle page can show what the people a keyword brought actually said, and set the terms your ads were clicked for beside the topics people raised once they arrived. The comparison is deterministic (shared significant tokens after stop-word removal) and auditable; where AI-written findings are added on top, a second model on a different provider reviews them and any sentence whose number is not in the computed figures is dropped before publication. Nothing is ever applied to your Google Ads account from this report.
Sending to GA4 needs no OAuth and nothing installed beyond the tag you already have: you paste a Measurement ID and an API secret (stored encrypted), and Karmaflow.ai's servers send events through Google's Measurement Protocol, to the global or the EU endpoint as you choose.
The direct path reports a conversion to your Google Ads account by click id, or by hashed email and phone (Enhanced Conversions for Leads) when there is no click id. It is a write, so it is gated four times, in order: a platform-wide switch, your workspace's approval for Google data sharing, the account's own switch, and a named conversion action of the accepting type. Uploads happen only from a workflow step you build; nothing uploads automatically.
The click id is looked for in order: the run itself, the contact, the contact's most recent chat session with a click id, a Google lead-form lead. Each upload carries an order id (by default the trigger and event id, or a field you map such as the deal id), so a re-run cannot double-count. Google answers per row, and each attempt is logged with its result — accepted, already recorded, rejected with Google's reason, or not sent with the gate that stopped it — for 90 days, with the click id stored only as a hash.
Where it is enabled for your workspace, one administrator's Google grant with the read-only Analytics scope lets Karmaflow.ai read aggregate reports from the property you send to: landing pages by channel, the paid-search queries GA4 credits, event counts by name and day, and a reconciliation of what we sent against what GA4 recorded. Reads are nightly snapshots with a probe of the property's own metadata first; nothing reads GA4 on a page view, and Google's refusals are stored beside each report.
The same tracker that captures identity is a full first-party analytics collector. Events are posted in batches to a collection endpoint (CORS, origin allow-list, bot filtering, IP hashed after geo lookup), streamed into a BigQuery events table partitioned by day and clustered by workspace, site and event name, and rolled up into per-workspace collections for dashboards. A tag manager compiles tags, triggers and variables into a versioned container served beside the tracker; one tag type fires events through your own Google tag on the page. The full description is in Web Analytics.
Every record below lives in your workspace's own database unless stated.
| Record | Holds | Retention |
|---|---|---|
| Google Ads integration | Account ids, link state, encrypted grant, time zone, currency, cursor, behaviour settings, quota counters, nightly snapshots (search terms, verification, budgets, change history) | Life of the connection |
| Lead record | Google's lead fields, consumer name and phone as supplied, typed identity, reply state, AI handling, erasure and billing stamps | Kept; erasure stamp on wipe-out |
| Message thread and messages | The conversation, attachments, who wrote each turn, approval state | Your transcript retention policy |
| Sent-message provenance | Every reply with provider, agent or user, thread, contact and workflow links | Kept |
| Lifecycle facts and lead outcomes | Per-day facts with tier and compute version; attributed outcomes with rung and evidence | Kept; re-synced over the trailing window |
| Ad clicks and creatives | Click id → campaign, ad group, keyword, ad; cached ad text | 455 days |
| Form leads | Lead-form submissions with campaign, ad group, ad and click id | Kept |
| Contact analytics identity | GA client id, session id, click ids, UTM, landing path, resolved ad | Life of the contact |
| GA4 integration | Measurement ID, encrypted API secret, region, toggles, event map | Life of the connection |
| GA4 delivery log | One row per send or skip, with reason; client id as a hash only | 30 days |
| Conversion uploads | One row per attempt with result and Google's reason; click id hashed | 90 days |
| GA4 report snapshots | Aggregate rows per report and property | Replaced nightly |