Drawing 02 · client anonymized

Immigration platform + IRCC autofiller

Bilingual intake for an immigration consultancy: a conditional assessment, structured client data, and a Chrome extension that fills the government portal.

Context

From November 2024 to August 2025 I worked with a regulated Canadian immigration consultancy: one RCIC-licensed consultant and their staff, serving Arabic-speaking applicants. A development team built the website and the admin dashboard. I wrote the specification that directed their work, designed the data model and the Free Assessment, and built the form-filling pipeline myself.

Problem

The last step of an application is IRCC’s online portal. For EMPP, its digital forms are IMM 0008, IMM 5406, IMM 5669 and IMM 5562, with pages repeated for the spouse and the children. Client details used to arrive in email chains, and staff transcribed them by hand into PDF forms and the portal.

  • Sensitive data. Passport and national ID numbers, ten years of addresses, criminality and health questions.
  • Two languages. Every client-facing string in English and Arabic.
  • IRCC owns the questions. The data has to map to the government’s fields one by one, and those fields change.
constraint
The consultant decides; the system prepares. The assessment says “you may be qualified” and books a consultation. The extension fills a page and stops. A person reviews it and submits.

What I built

The Free Assessment. I designed it as a bilingual decision tree; on the site it runs as a multi-step form. It collects contact details, then branches. Applying from inside Canada leads to eight services, from Super Visa to PR renewal. Otherwise, questions on work, family, IELTS scores, funds and refugee status lead toward Express Entry, OINP, LMIA, EMPP, private sponsorship (SAH), or a visitor, business or study visa. The tree encodes concrete thresholds: IELTS 8/7/7/7 with a two-year diploma, a bank balance above USD 20,000, IRCC’s proof-of-funds table by family size. Every result ends in a Calendly booking link, except one: “no programs available at this time”. Answers are stored through the API.

An applicant answers a bilingual Free Assessment. In this excerpt, applying from inside Canada leads to eight in-Canada services. Other paths pass through skipped questions to a language-and-funds check that leads to Express Entry or OINP, or to a UNHCR or UNRWA registration check that leads to EMPP. Every outcome reaches the consultant dashboard. The client record, holding up to 18 people, flows through the IRCC form field map to the Chrome extension, which fills the IRCC portal for staff to review and submit.applicantEN · ARin Canada?yes8 in-CanadaservicesnoIELTS + funds?UNHCR/UNRWA?yesyesExpress Entryor OINPEMPPconsultant dashboardaccounts · form packagesclient record18 people per intakeIRCC form field mapLLM-extracted · JSONLChrome extensionfills by field nameIRCC portalstaff review, submit
Fig. 1 An excerpt of the assessment (dashed lines skip questions), then the designed data path. In v1 the extension read hand-loaded JSONL files, not the client record.

A specification and a data model. I wrote the platform spec in five parts, plus a database design. The model mirrors the forms. Each IRCC question becomes one short key: “Passport number” is passport_num, “Date of birth (YYYY-MM-DD)” is birth_date. One person table holds everyone on an application, tagged by relationship. Addresses, education, travel and criminality answers hang off each person. The intake covers 18 people: the principal applicant and the spouse, each with up to three children, two parents and up to three siblings.

In the spec, staff assign a package, such as EMPP, and the client portal shows only those forms. For the review side I prototyped a data viewer: one card per person, a modal with everything on file, and click-to-copy on every value. The spec keeps it read-only; only an admin can toggle editing.

The spec routes static site text through a Google Sheet: 109 keys in 11 sections, English and Arabic side by side. A developer runs a script that turns it into en.json and ar.json, so copy changes need no code.

An LLM pipeline over the government’s forms. I saved a complete local copy of the portal’s forms: 65 pages: the portal’s two main pages and the four IMM forms. A Python script stripped each page down to its form markup, dropping scripts, styles and session-timeout dialogs. An LLM read the 60 stripped form pages and returned every field as structured data: label, name and id, type, options, required flag, help text, and repeating groups like children or trips. A second script kept the principal applicant’s 217 text fields as JSONL, one line per field: section, key, label, value. Family members’ pages reuse most of the principal applicant’s field names (190 of the spouse’s 231), so the same format works for each person. The model did the reading; every key is the portal’s own name attribute.

65IRCC portal pages captured
823field records extracted
217text fields in the JSONL

A Chrome extension. My first plan proposed a Tampermonkey userscript. v1, packaged in June 2025, is a Manifest V3 Chrome extension for staff. Load a person’s JSONL file, open a portal page, press Fill Form. It finds each field by name, writes the value, and fires input and change events so the portal’s form code registers it. Then it shows “Please review the form” and stops. There is no submit step.

How I knew it worked

Offline first. The local copy let me build and test without the live portal. Each page works on its own; only the links between pages are dead. Three sample applicant files exercised the fill.

Real keys. All 170 distinct keys in the JSONL appear as name attributes in the 60 stripped pages. The model transcribed them and invented none.

Status at handover. In August 2025 I handed over the spec, the scripts, the sample data, an issue list and a 36-minute walkthrough video. The spec records where each part stood. The Free Assessment worked and needed no changes. The services and news CMS was complete. The extension was a functional proof of concept. The client portal was still in development.

failure

v1 was delivered with four known bugs, all in the issue list. Passport and national ID dates got mixed up. So did everyone’s education dates. The spouse’s trip destination wouldn’t fill. Dates in general were unreliable.

The first two share one cause, and I built it in: fields keyed by the name attribute alone. My plan called name stable and “unique per field”. It is stable, not unique. The passport and national ID pages both use issueDate and expiryDate. Five IMM 5669 pages, from education to addresses, all use from0 and to0. In the JSONL, 27 keys repeat across 74 of the 217 rows. The fill loop writes every row whose name matches, so the last one wins.

v1 also can’t pull from the platform. Each person’s data has to be prepared as a JSONL file and loaded by hand. That gap is why the v2 design exists: staff sign in with dashboard credentials, pick an application from the API, and fill from live data. It is designed, not built.

What I’d change

Make keys unique, and check it. Form, page and name, not name alone. The converter should refuse a duplicate key instead of letting it surface as a wrong date in the portal.

Check the extraction before trusting it. The model found a label for every field except dates: all 64 date fields came back “Label not found”. A reviewer can’t verify a value they can’t identify. I’d treat an unlabelled field as an extraction failure, not as data.

Carry every field type. The converter kept text inputs only: 217 of the 284 fields on the principal applicant’s pages. The fill code already handled selects, radios and checkboxes; the data files never carried them.

None public. The client is anonymized and the platform is private.