Datasite
Datasite v1.2.0
Publisher description
From the marketplace listing
Connect ChatGPT to your Datasite virtual data room - the secure workspace where thousands of M&A deals are facilitated annually. Set up folder structures, invite users, search documents, track buyer Q&A, and audit data room readiness, all through natural language. No workflow interruptions. No security trade-offs. Built for advisors, bankers, and corporate development teams - backed by the enterprise security and permissioning every transaction demands. 1. Set up a deal room - "Create a new Datasite Prepare data room for Project Alpha, set up a standard M&A folder index, and invite sarah@advisorfirm.com as an admin." 2. Search documents semantically - "Search the data room for any documents related to pending litigation or regulatory risk." 3. Track buyer Q&A - "Show me all open Q&A questions in the data room and draft responses where we have supporting documents." 4. Audit deal room readiness - "Check the data room for empty folders, missing documents, and any files flagged as password-protected or blank before we go live." 5. Manage user access - "Create a buyer role with view-only permissions and create draft invitations for the following five contacts from Acme Capital."
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Skill instructions
bulk-qa-answers14.5 KB
--- name: bulk-qa-answers description: > Bulk Q&A Answers skill for Datasite deal rooms. Use this skill whenever a sell-side deal team wants to answer multiple buyer questions at once, generate AI draft responses from VDR content, produce a Q&A tracker spreadsheet, or build a Q&A management dashboard. Triggers include: "answer the Q&A", "draft responses to buyer questions", "process the question list", "generate Q&A tracker", "answer all questions", "bulk answer", "Q&A management dashboard", "respond to diligence questions", or any request to systematically work through a list of buyer questions using data room content as the source. Use this skill proactively whenever a buyer has submitted questions and the deal team wants AI-assisted drafting. Do not use for individual one-off questions outside a structured Q&A process. Do not draft answers from general knowledge — all responses must come from the data room. metadata: author: Blueflame AI version: 1.0.0 mcp-server: datasite category: deal-management tags: [datasite, vdr, m&a, q-and-a, due-diligence, blueflame] --- # Bulk Q&A Answers You are helping a sell-side deal team draft answers to buyer due diligence questions by reading and interpreting Datasite data room content. You produce two outputs: a formatted Excel tracker and an interactive React Q&A management dashboard. --- ## Terminology — fileroom vs. folder Use these terms precisely when communicating with the user: - **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas. - **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms. When in doubt: if it is not the single top-level container for the whole project, it is a folder. ## Feature Requirements | Capability | Free | Requires Blueflame | |---|:---:|:---:| | Q&A status overview and health metrics | ✅ | — | | Draft answers to buyer questions from document content | — | ✅ | | Source citations with document name, path, and page number | — | ✅ | | Excel tracker and dashboard | ✅ | — | **Without Blueflame:** The skill can retrieve the Q&A status overview and display question counts and categories. It cannot draft answers — all questions will be marked Open. The core value of this skill requires Blueflame. **With Blueflame:** `searchDocuments` finds relevant passages in the data room for each question and drafts a professional sell-side response grounded in document content, with full source citations. > ⚠️ **Blueflame content guard — mandatory** > `searchDocuments` is the only permitted source of document content. > - **Never draft Q&A answers from Claude's general knowledge.** All responses must be grounded in data room documents retrieved via `searchDocuments`. A fabricated answer is worse than no answer. > - If `searchDocuments` returns an **activation link** instead of results, **stop immediately** and tell the user: > > > "To draft answers grounded in your data room, Blueflame AI search needs to be activated on this project: > > 🔗 **Activate Blueflame:** [activation link] > > **With Blueflame:** I'll read the relevant documents for each question and draft a professional sell-side response citing the document name and page — so every answer is defensible and traceable back to source. Without it I have no way to read what's in your data room and cannot draft responses. > > Please activate Blueflame and then re-run." > > - Do not attempt to draft any answers until content search is confirmed working. > - All Q&A answers **must** be sourced exclusively from tool results. ## Step 1 — Load the questions The user will provide a spreadsheet of questions. Read it and extract for each row: - Question text - Buyer group / individual who asked it (the "Question From") - Any existing status, category, or section grouping already in the file - Any prior answer already provided (skip these unless the user asks to re-draft) If any column mappings are unclear, ask the user to confirm before proceeding. --- ## Step 2 — Understand the deal context Call `getProjectOverview` to confirm the project name, sector, and fileroom structure. This orients your research — you'll know which areas of the data room are likely relevant for each question type (e.g. financial questions → Finance folder, IP questions → Technology/IP folder). --- ## Step 3 — Research and draft each answer For each unanswered question, use the following research workflow. The goal is not just to locate a document but to **read and interpret its content** so the answer reflects genuine understanding of the material. ### 3a — Semantic search first (primary) Run `searchDocuments` with the question (or a distilled version of it) as the query. Use `decompose: true` for complex or multi-part questions — this breaks the query into sub-queries and finds relevant passages across the whole data room that keyword search would miss. `searchDocuments` returns text passages with document names, page numbers, and relevance scores. Read the passages — they are actual document content, not just file names. Use them to understand what the data room says on the topic. ### 3b — Keyword search for specifics (secondary) After the semantic search, run `searchDocuments` for any specific terms, figures, or exact phrases that the question calls for — e.g. a specific contract name, a company name, a regulation, a year, a metric. Keyword search complements semantic search for precise lookups. ### 3c — Browse to the relevant folder if needed If the search results point to a specific section of the data room but you need to confirm what documents are present (e.g. to note which years of accounts are filed, or whether a specific agreement exists), use `listFolderContents` to navigate to that folder and inspect its contents directly. ### 3d — Synthesise and draft the answer With the passages and document context in hand, write a clear, factual response. The standard to aim for: - **Directly answers** what was asked — not a broader essay on the topic - **Grounded in the documents** — reflects what the data room actually says, not general knowledge - **Sell-side voice** — professional, concise, confident. Written as if the CFO or GC reviewed it, not as a transcript of search results - **Handles uncertainty correctly** — if the data room contains partial information, say so clearly (e.g. "Management accounts for FY2024 and FY2025 are available; audited accounts for FY2023 are not yet uploaded"). Never fill gaps with assumptions. - **Sensitive matters** — if a question touches on active litigation strategy, unpublished projections, or personal employee data, flag it for legal review rather than drafting a response ### 3e — Assign a status - **Complete** — question fully answered with clear source material - **Partial** — answer drafted but source material is incomplete or only partially responsive - **Open** — insufficient source material found; needs manual input from the deal team ### 3f — Build the source reference and citation For every answer, record two things: **Source Reference** (brief, for the tracker): the VDR folder path and document name — e.g. `3.1 Audited Accounts / FY2024 Annual Report` or `5.3 Customer Contracts / MSA with Acme Corp` **Document Citation** (detailed, for verification): the full citation including document name, VDR index path, and page number(s) where the relevant content was found — e.g. `FY2024 Annual Report (VDR 3.1), p.14 — Revenue recognition policy` or `Employment Agreement — J. Smith (VDR 7.2.4), p.3 — Clause 8, Non-compete`. If multiple documents were used, list each on a separate line. If no source is found after running both semantic and keyword searches and browsing the relevant folder, mark the question Open and note: "No source material found in data room — requires manual response." --- ## Step 4 — Group questions by theme Before producing outputs, group questions into thematic sections. Common M&A Q&A groupings: - Financial Performance & Accounting - Tax - Legal & Regulatory - Commercial & Customers - Human Resources & Management - Intellectual Property & Technology - Operations - ESG & Environmental - Other / Miscellaneous Use the question content (and any category column already in the input file) to assign each question to a section. --- ## Step 5 — Offer outputs Before generating the Excel tracker and dashboard, ask: > "I've drafted answers for all [N] questions. What would you like me to produce? > - **Excel tracker** — formatted spreadsheet with all questions, answers, statuses, and source citations > - **Q&A management dashboard** — interactive React dashboard for active deal management (uses additional credits) > - **Both** > - **Neither** — just show me the answers in this conversation" Only generate the Excel tracker and/or dashboard if the user explicitly requests them. ## Step 5b — Produce the Excel tracker (only if requested) Use the xlsx skill to produce a formatted `.xlsx` file saved to the outputs folder. **Columns (in order):** 1. **Diligence Question** — the original question text verbatim 2. **Diligence Response** — the AI-drafted answer 3. **Status** — Complete / Partial / Open 4. **Question From** — buyer group or individual name 5. **Source Reference** — VDR folder path and document name (brief) 6. **Document Citation** — full citation with document name, VDR index, page number(s) and clause/section where relevant. Multiple sources listed on separate lines within the cell. **Formatting rules:** - Header row: dark blue background (`#1a2332`), white font, bold - For each new theme/section, insert a **separator row** spanning all 6 columns containing the section name, styled with mid-blue background (`#2d4a6e`), white bold text — a visual divider, not a data row - Enable **text wrapping** on the "Diligence Response" column (column B) and "Document Citation" column (column F). Set column widths: B ~60 chars, F ~50 chars - Status cell colour coding: Complete = light green fill, Partial = light amber fill, Open = light red fill - Freeze the header row Save as `[ProjectName]_QA_Tracker_[Date].xlsx` in the outputs folder. --- ## Step 6 — Produce the React Q&A management dashboard (only if requested) Read `references/dashboard-spec.md` for the full React component specification before building. The dashboard is a self-contained React component populated with the actual questions, answers, statuses, buyer groups, source references, and citations generated during the Q&A drafting process. It is for active deal management — it should feel live and usable, not like a static report. Key sections to implement (details in the reference file): 1. **Summary KPI Bar** — four stat cards (Total, Open, Awaiting Review, Submitted) 2. **Past Q&A Trackers** — collapsible card with drag-and-drop upload zone for precedent deals 3. **AI Buyer Group Q&A Analysis** — collapsible panel with per-buyer stats, topic volume charts, and AI strategic signal 4. **Filter Bar + Question Log** — searchable, filterable list with expandable rows showing AI draft, VDR citations, and management feedback thread 5. **Dashboard Modal** — full KPI and analytics view with time-savings metrics Use navy `#1a2332` / gold `#d4a017` colour palette with Source Sans 3 font. All state via `useState` — no backend required. --- ## Step 7 — Deliver to the user Present both outputs: 1. Link to the Excel tracker file 2. The React dashboard artifact rendered in the conversation Then say: > "I've drafted answers to [N] questions — [X] Complete, [Y] Partial, [Z] Open. The [Z] open questions need manual input as I couldn't find sufficient source material in the data room. Both the Excel tracker and the live dashboard are ready above." If there are Partial answers, offer: > "For the [Y] partial answers, want me to flag the specific gaps so the team knows exactly what additional material to source?" --- ## Operating principles **Read the documents, don't just locate them.** `searchDocuments` returns actual text passages — use them. The quality of the answer depends on understanding what the document says, not just knowing it exists. **Source everything.** Every drafted answer must have a citation. Unsourced answers should be marked Open. Buyers will scrutinise these responses — a wrong answer is worse than no answer. **Write in the seller's voice.** Concise, factual, professional. Not a summary of search results. **Don't over-answer.** Answer the specific question asked. Buyers will follow up for more. **Flag patterns.** If multiple buyers ask the same question, note it — it signals an IM gap or a known concern the deal team should address proactively. **Respect sensitivity.** Active litigation strategy, unpublished projections, and personal employee data should be flagged for legal review, not drafted. ## Performance Notes - **Quality over speed.** A wrong answer is worse than no answer — buyers will scrutinise every response. - Read the source passages returned by `searchDocuments` fully before drafting. Do not skim. - Do not skip the keyword search step for questions involving specific figures, dates, or names. - Mark questions Open rather than guessing when source material is insufficient. --- ## Common Issues **`getProjectOverview` fails or returns the wrong project** Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use. **`listFolderContents` returns no results** The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available. **`searchDocuments` returns an activation link instead of results** Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge. **MCP disconnects mid-workflow** Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched. **`updateContent` or `createContent` returns a permissions error** The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
document-quality-check27.1 KB
---
name: document-quality-check
description: >
Document Quality Check skill for Datasite deal rooms. Use this skill whenever a
deal team wants to audit document quality before going live to buyers. Triggers
include: "check document quality", "flag bad documents", "find password protected
files", "check for blank documents", "PII check", "redaction review", "find
corrupted files", "document audit", "quality check the data room", "are there
any blank or broken files", "check for unredacted personal data", or any request
to verify that documents in the data room are complete, accessible, and safe to
share. Use this skill proactively before a data room goes live.
Do not use for renaming files (use smart-file-renaming) or for identifying
missing sections (use gap-analysis).
metadata:
author: Blueflame AI
version: 1.0.0
mcp-server: datasite
category: deal-management
tags: [datasite, vdr, m&a, document-quality, pii, redaction, blueflame]
---
# Document Quality Check
You are helping a deal team verify that every document in their Datasite data room is fit to share with buyers before going live. You check for six categories of quality issues and produce an HTML dashboard with a downloadable Excel report.
---
## Terminology — fileroom vs. folder
Use these terms precisely when communicating with the user:
- **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas.
- **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms.
When in doubt: if it is not the single top-level container for the whole project, it is a folder.
## Feature Requirements
| Capability | Free | Requires Blueflame |
|---|:---:|:---:|
| Failed / unprocessable files (status metadata) | ✅ | — |
| Placeholder and stub documents (type metadata) | ✅ | — |
| Wrong file formats (fileType metadata) | ✅ | — |
| Uninformative filenames (name pattern matching) | ✅ | — |
| Version conflicts (name pattern matching) | ✅ | — |
| Duplicate documents (name + metadata comparison) | ✅ | — |
| Stale documents (upload date metadata) | ✅ | — |
| PII exposure in document content | — | ✅ |
| Redaction quality check | — | ✅ |
| Broken references and missing exhibits | — | ✅ |
**Without Blueflame:** 7 of 10 checks run fully using `listFolderContents` metadata. The three content-level checks (PII, redaction quality, broken references) are skipped — note these in the report as "Requires Blueflame."
**With Blueflame:** All 10 checks run. `searchDocuments` scans document content for PII patterns, verifies redaction quality, and finds broken cross-references inside documents.
> ⚠️ **Blueflame fallback — two-tier behaviour**
> `searchDocuments` is the only permitted source of document content.
> - Do **not** infer document content, PII presence, or redaction quality from Claude’s training knowledge.
> - **Phase A** (Checks 1–7, metadata checks) uses `listFolderContents` only — always free. Complete Phase A fully first.
> - **Phase B** (Checks 8–10: PII scan, redaction quality, broken references) requires `searchDocuments`. When you reach Phase B, attempt one call. If it returns an **activation link** instead of results:
> 1. **Do not generate the HTML dashboard yet** — ask the Blueflame question first as a plain conversational message
> 2. Summarise Phase A findings in plain text (e.g. "I found X password-protected files, Y duplicates, Z files with no extension")
> 3. Then ask:
>
> > "I’ve completed the 7 metadata checks — here’s what I found: [plain text summary]. To also run PII scanning, redaction quality checks, and broken reference detection, Blueflame AI search needs to be activated on this project:
> > 🔗 **Activate Blueflame:** [activation link]
> > **With Blueflame:** I’ll scan document content for exposed personal data (names, NI numbers, bank details), verify that redacted text can’t be read in the file layer, and check for broken cross-references inside documents — the checks buyers and their lawyers look for most.
> > Would you like to activate now, or shall I produce the dashboard with the Phase A findings only?"
>
> 4. **Wait for the user’s response before producing any dashboard or output file.**
>
> Do not embed the Blueflame activation prompt inside the HTML dashboard — it must appear as an interactive conversational question before any output is generated.
> **`listFolderContents` — efficient traversal**
> - `depth: 1` (default) — immediate children only. Use for targeted lookups.
> - `depth: 5, foldersOnly: true` (default when depth > 1) — full folder tree in one call, no documents. Use for structural checks.
> - `depth: 5, foldersOnly: false` — full folder tree including all document metadata in one call. Use when building a document inventory.
> - When `depth > 1`, the response is a **flat list** with `depth` and `path` columns — not a nested tree.
## Step 1 — Orient yourself
Call `getProjectOverview` to understand the project structure and get a list of all filerooms. You'll work through each fileroom systematically.
---
## Step 2 — Run all quality checks
Work through each check below in two phases:
**Phase A — Metadata checks (use `listFolderContents`):**
Call `listFolderContents` without a metadataId to get all filerooms, then recurse into each folder to build a complete document inventory. Each document entry includes: name, type (DOCUMENT/PLACEHOLDER/FOLDER/FILE_ROOM/SANDBOX), status (DONE/FAILED/PROCESSING), fileType (pdf/docx/xlsx etc.), publishingState, and upload date. Use this single inventory pass to run all metadata-based checks — do not make a separate call per check.
From the inventory, flag:
- `status: FAILED` or `PROCESSING` → Check 1 (unprocessable)
- `type: PLACEHOLDER` → Check 6 (placeholder/stub)
- `fileType` in [xlsm, xlsb, zip, rar, msg, eml, pages, numbers, key, dwg] → Check 9 (wrong format)
- Name patterns: Scan/IMG/Document/Untitled/Copy of/FINAL_FINAL/USE THIS/DO NOT USE → Check 12 (bad filenames)
- Name patterns: v1/v2/revised/updated/superseded/old/archive → Check 8 (version conflicts)
- Identical names in the same folder → Check 7 (duplicates)
- Upload date > 12 months ago in active sections (Management Accounts, Insurance, Licences) → Check 10 (stale)
**Phase B — Content checks (use `searchDocuments`):**
Use `searchDocuments` for checks requiring reading inside documents (PII, redaction quality, broken references). Always call `searchDocuments` — if AI search is not yet activated the tool returns an activation link; present it to the user rather than skipping the check.
---
### Check 1 — Failed / unprocessable documents
*Catches: password-protected files, corrupted files, files that couldn't be indexed*
```
listFolderContents(projectId, query="*", filter=["status:EQ:FAILED", "type:EQ:DOCUMENT"])
listFolderContents(projectId, query="*", filter=["status:EQ:PROCESSING", "type:EQ:DOCUMENT"])
```
`FAILED` = Datasite could not process the file — most commonly because it is password-protected or corrupted. `PROCESSING` documents that have been in that state for more than a few minutes are likely stuck (possible corruption or unsupported format).
For each result note: filename, VDR folder path, file size, extension.
**Severity:** High — buyers cannot open these documents.
---
### Check 2 — Blank or near-blank documents
*Catches: accidentally uploaded blank pages, empty documents, placeholder files*
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "pageCount:LT:2", "fileSize:LT:50000"])
```
A document with fewer than 2 pages AND under 50KB is almost certainly blank or a single near-empty page. Cross-reference against the folder context — a 1-page certificate of incorporation is fine; a 1-page "FY2024 Audited Accounts" is not.
Also flag zero-byte files:
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "fileSize:LT:1000"])
```
For each result, check the filename and folder path to judge whether the low page count is expected. Flag only where it looks wrong for the document type.
**Severity:** High (if it's a material document), Medium (if it's a supporting file).
---
### Check 3 — Suspicious redaction quality (poorly blacklined documents)
*Catches: documents with redactions that may be incomplete or incorrectly applied*
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "redacted:EQ:true"])
```
This returns all documents Datasite has flagged as containing redactions. For each one, use `searchDocuments` to check whether sensitive content that should have been redacted is still readable — e.g. if a document is flagged as redacted but the underlying text was not properly removed (a common issue with image-based PDFs where black boxes are overlaid on text that remains in the file layer).
Search queries to run on redacted documents:
- `searchDocuments` with query "salary" or "compensation" — check for unredacted pay figures
- `searchDocuments` with query "date of birth" or "national insurance" — check for unredacted personal identifiers
- `searchDocuments` with query "account number" or "IBAN" — check for unredacted banking details
Flag any redacted document where searchable text appears beneath the redaction, or where the expected content is still visible in snippets.
Also flag documents where the filename suggests redaction was needed (e.g. "Employment Agreements", "Payroll", "Personal Data") but `redacted:EQ:false` — these may have been shared without any redaction applied.
**Severity:** High — unredacted personal or sensitive data in a buyer-facing data room is a GDPR/privacy breach.
---
### Check 4 — PII exposed without redaction
*Catches: personal data visible in documents that haven't been redacted at all*
Run the following `searchDocuments` queries across the full data room. Each targets a specific PII category. Read the snippets returned and flag any document where personal data is clearly visible.
**Personal identifiers:**
- `"date of birth"` or `"DOB"` or `"born on"` — personal birth dates
- `"passport number"` or `"passport no"` — passport identifiers
- `"national insurance"` or `"NI number"` or `"social security"` or `"SSN"` — government ID numbers
- `"home address"` or `"residential address"` — personal addresses (distinguish from business addresses)
- `"driving licence"` or `"driver's license number"` — licence identifiers
**Financial details:**
- `"sort code"` and `"account number"` — personal bank account details
- `"IBAN"` — international bank account numbers
- `"salary"` with a named individual — personal salary data linked to a person's name
- `"payslip"` or `"pay stub"` — payroll documents that typically contain personal financial data
**Contact data:**
- Search for personal email domain patterns: `"@gmail.com"` or `"@yahoo.com"` or `"@hotmail.com"` or `"@icloud.com"` — personal email addresses (business emails like @companyname.com are expected and fine)
- `"mobile"` or `"personal phone"` alongside a person's name — personal phone numbers
For each snippet returned, assess whether it appears in a context that warrants redaction (e.g. an employee's salary in a payroll schedule = High risk; a reference to "date of birth required for background check" in an HR policy = Low risk).
**Severity:** High for direct identifiers (passport, NI/SSN, bank account); Medium for contact data and salary where it's incidental.
---
### Check 5 — Suspiciously small or potentially incomplete scanned documents
*Catches: multi-page documents where pages may be missing*
For scanned documents (PDFs from physical paper), page count alone can reveal gaps. Use:
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:pdf"], sort=["pageCount,ASC"])
```
Cross-reference page count against what's expected based on document type and filename:
- A "Lease Agreement" with 2 pages is suspicious — commercial leases are typically 20–100 pages
- An "Employment Agreement" with 1 page is suspicious — these typically run 5–30 pages
- An "Audited Financial Statements" document with 3 pages is suspicious — audited accounts are typically 30–100+ pages
- A "Certificate of Incorporation" with 1–2 pages is fine
Flag documents where the page count appears materially below what the document type would normally require. Note the filename, VDR path, current page count, and the expected range.
**Severity:** Medium — missing pages may mean incomplete disclosure. High if it's a key legal or financial document.
---
### Check 6 — Placeholder or stub documents
*Catches: files named as placeholders, zero-content uploads, "TBC" files*
```
listFolderContents(projectId, query="placeholder", filter=["type:EQ:DOCUMENT"])
listFolderContents(projectId, query="TBC", filter=["type:EQ:DOCUMENT"])
listFolderContents(projectId, query="draft", filter=["type:EQ:DOCUMENT"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:placeholder"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:TBC"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:WIP"])
```
Also check for documents with generic names that suggest they haven't been properly named or are still in progress: "Document1", "Untitled", "Copy of", "v1", "DRAFT", "temp".
**Severity:** Medium — placeholder documents signal incomplete preparation; buyers will notice.
---
### Check 7 — Duplicate documents
*Catches: exact or near-duplicate files that inflate apparent completeness and expose version inconsistencies*
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT"], sort=["fileSize,ASC"])
```
Group results by file size. Where two or more documents share the same `fileSize`, compare their filenames. Exact file size match + near-identical filename = likely duplicate. Also flag same `pageCount` + same folder path with minor filename variation (e.g. `Agreement_v1.pdf` and `Agreement_final.pdf` in the same folder).
For suspected duplicates in high-risk areas (financial statements, contracts), run `searchDocuments` on both documents to compare leading paragraphs — if content is near-identical, flag as a confirmed duplicate.
Highest risk: duplicate financial statements or contracts where versions may differ in a key figure or clause.
**Severity:** High (if material documents like contracts or financials are duplicated with differing content), Medium (identical duplicates — one just needs removing).
---
### Check 8 — Version conflicts and superseded documents
*Catches: old or draft versions left in the room alongside current ones, which buyers may read and draw incorrect conclusions from*
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:v1"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:revised"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:updated"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:superseded"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:previous"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:archive"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:old"])
```
Also search for: `"v2"`, `"final"`, `"draft"` in filenames. Where multiple versions exist in the same folder, flag all but the most recently modified (`sort: availableDate,DESC`) as potentially superseded.
Cross-reference `availableDate` against document content date where visible — a file uploaded in 2026 but containing a 2023 date header warrants flagging.
**Severity:** High (if two versions of a contract or financial statement coexist with potentially different terms or figures), Medium (clear drafts or superseded copies that are obviously not current).
---
### Check 9 — Wrong file format or rendering risk
*Catches: files that buyers cannot open in-browser, macro-enabled files (security risk), and archive files that block search indexing*
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:msg"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:eml"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:xlsm"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:xlsb"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:zip"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:rar"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:dwg"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:pages"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:numbers"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "extension:EQ:key"])
```
Flag each by issue type:
- `.xlsm` / `.xlsb` → macro-enabled Excel — security risk for buyers, may be blocked by corporate IT; recommend saving as `.xlsx`
- `.zip` / `.rar` → archive files — content invisible to VDR search indexing, buyers cannot open in-browser; recommend unpacking and uploading individual files
- `.msg` / `.eml` → email files — rarely intentional, likely contain unintended PII or privileged content; recommend converting to PDF
- `.dwg` / `.dxf` → CAD files — buyers without AutoCAD cannot open; recommend PDF export
- `.pages` / `.numbers` / `.key` → Apple-native formats — Windows users (most buyers) cannot open; recommend PDF or Office format
**Severity:** High for `.msg`/`.eml` (PII/privilege risk) and `.zip` (invisible to search), Medium for rendering-incompatible formats.
---
### Check 10 — Stale or outdated documents
*Catches: documents that appear current but haven't been updated in over a year, particularly in areas where buyers expect current data*
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "availableDate:LT:[12_months_ago_epoch]"], sort=["availableDate,ASC"])
```
Calculate the epoch timestamp for 12 months ago from today's date and substitute it into the filter. From the results, focus on document types where staleness is a material risk:
- Management accounts — must be current; flagging anything older than 3 months
- Financial models and projections — flag if older than 6 months
- Employee lists and org charts — flag if older than 12 months
- Insurance schedules — flag if upload date is older than 12 months (policy may have expired)
- Regulatory licences and certificates — flag if older than 12 months (renewal may be overdue)
- Board minutes — flag if the most recent entry is older than 6 months
Don't flag inherently historical documents (e.g. FY2022 audited accounts — they're supposed to be from 2022).
**Severity:** High (management accounts, insurance, regulatory licences past renewal date), Medium (financial models, employee lists).
---
### Check 11 — Broken references and missing linked content
*Catches: documents referencing exhibits, appendices, or schedules that were never uploaded — buyers encounter dead ends*
Use `searchDocuments` with the following queries:
- `"see attached"` or `"refer to appendix"` or `"as per schedule"` — cross-reference whether the referenced exhibit exists in the same folder
- `"exhibit"` or `"annex"` or `"schedule"` — check whether named attachments are present
- `"[TBC]"` or `"[insert"` or `"[link]"` or `"[see tab"` — internal authoring placeholders never resolved before upload
- `"see accompanying"` or `"as set out in"` or `"detailed in the attached"` — general cross-reference language
For each match, check whether the referenced document is present in the same folder using `listFolderContents`. Flag where it is absent.
For Excel financial models specifically: if a CIM or management presentation references a "detailed financial model" and the only Excel file in the folder has `fileSize:LT:100000`, it is likely a stub or broken-link version — flag for review.
**Severity:** High (missing exhibit to a contract, missing appendix to audited accounts), Medium (unresolved placeholder text).
---
### Check 12 — Uninformative or unprofessional filenames
*Catches: filenames that signal poor preparation and make navigation impossible for buyers — a direct reputational risk*
```
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:Scan"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:Document"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:Copy of"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:FINAL_FINAL"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:USE THIS"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:DO NOT USE"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:Untitled"])
listFolderContents(projectId, query="*", filter=["type:EQ:DOCUMENT", "name:LIKE:New "])
```
Also flag:
- Files with sequential numbering in the name: match patterns like `001`, `002`, `(1)`, `(2)`, `(3)`
- Files with double extensions: `.pdf.pdf`, `.docx.pdf` — artefacts of bulk upload tools
- Any folder where more than 20% of filenames match these generic patterns — flag the entire folder for a renaming pass, not just individual files
**Severity:** Medium across the board — these don't block access but signal poor preparation to buyers. Flag the folder-level pattern as more severe than individual files.
---
## Step 3 — Compile findings
Compile all findings into a structured list:
```
findings = [
{
check: "Failed / Unprocessable",
severity: "High",
filename: "FY2024 Audited Accounts.pdf",
folder: "3.1 Audited Financial Statements",
detail: "Document status is FAILED — likely password-protected or corrupted. Buyers cannot open it.",
recommended_action: "Remove password protection or re-export as an unprotected PDF and re-upload."
},
...
]
```
Count issues by severity and check type for the dashboard scorecard.
---
## Step 4 — Offer the dashboard
Before generating anything, ask:
> "I've completed the quality checks. Would you like me to produce the HTML dashboard with the full findings and an Excel export? It uses additional credits to render. Alternatively I can give you a plain text summary here."
Only generate the dashboard if the user confirms. If they decline, go to Step 5 and deliver a plain text summary.
## Step 4b — Produce the HTML dashboard (only if requested)
Generate a self-contained HTML artifact. Include a **"Download as Excel"** button using SheetJS (`https://cdnjs.cloudflare.com/ajax/libs/xlsx/0.18.5/xlsx.full.min.js`) that exports the findings table with columns: Check Type | Severity | Filename | VDR Folder | Detail | Recommended Action | Status (default: Open).
**Dashboard structure:**
**Header:**
- Deal name, date of audit, total issue counts by severity (High / Medium / Low)
- "Download as Excel" button (navy, top right)
**Summary scorecard — six check tiles:**
One tile per check type, each showing:
- Check name and icon
- Issue count
- RAG status: Red (any High issue), Amber (Medium only), Green (no issues)
Check tiles:
- 🔒 Failed / Unprocessable
- 📄 Blank / Near-blank
- ✂️ Redaction Quality
- 👤 PII Exposed
- 📑 Incomplete Scans
- 📝 Placeholders / Stubs
- 👯 Duplicate Documents
- 🔁 Version Conflicts
- ⚠️ Wrong Format / Rendering Risk
- 🕐 Stale / Outdated Documents
- 🔗 Broken References
- 🏷️ Uninformative Filenames
**Findings table (below scorecard):**
- Filterable by check type and severity
- Columns: Severity badge | Check Type | Filename (with VDR folder path below in grey) | Issue Detail | Recommended Action
- Severity badges: High = red (`#ef4444`), Medium = amber (`#d97706`), Low = grey (`#6B7280`)
- Rows sorted High → Medium → Low within each check type
**Design:** white background, navy (`#1a2332`) header, 12px border-radius cards, `Source Sans 3` font via Google Fonts, no external dependencies beyond SheetJS and fonts.
---
## Step 5 — Deliver to the user
Give a brief summary:
> "I've checked [N] documents across [M] filerooms and found [X] High and [Y] Medium quality issues. The most urgent: [top 2–3 findings]. Use the Download button to export the full report as Excel for the team to action."
Then offer:
> "Want me to flag which issues are quickest to fix vs. which need the document owner involved?"
---
## Operating principles
**Context matters for severity.** A 1-page PDF is fine for a certificate; it's a red flag for an audited accounts file. Always check the filename and folder path before flagging a low page count.
**Don't cry wolf on PII.** A business email address in a contract is expected. A director's personal gmail address in a board minute is a flag. Read the snippet context before raising an issue.
**Redaction quality is a GDPR risk, not just a tidiness issue.** Documents where text is visually blocked but remains machine-readable in the PDF layer are the most dangerous scenario — prioritise these.
**Failed documents are the most urgent fix.** A buyer who clicks a document and gets an error immediately loses confidence in the deal team's preparation. Every failed document should be actioned before go-live.
**Be specific in recommended actions.** "Remove password protection and re-upload" is useful. "Fix document" is not.
## Performance Notes
- **Do not skip checks to save time.** A missed password-protected file or undetected PII exposure is a serious issue that could delay go-live or create a compliance breach.
- Run all Phase A checks from a single `listFolderContents` pass — avoid repeated calls.
- Context matters before flagging: always check filename and folder path before raising a severity issue.
---
## Common Issues
**`getProjectOverview` fails or returns the wrong project**
Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use.
**`listFolderContents` returns no results**
The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available.
**`searchDocuments` returns an activation link instead of results**
Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge.
**MCP disconnects mid-workflow**
Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched.
**`updateContent` or `createContent` returns a permissions error**
The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
gap-analysis18.7 KB
---
name: gap-analysis
description: >
Data Room Gap Analysis skill for Datasite deal rooms. Use this skill whenever a
sell-side deal team wants to audit what is missing, sparse, or incomplete in their
data room before going live to buyers. Triggers include: "run a gap analysis",
"what's missing from the data room", "check the data room coverage", "flag empty
folders", "what haven't we uploaded yet", "data room readiness check", "find gaps
before we go live", "are all the contracts in there", "check we have everything",
or any request to assess completeness of the data room by section. Use this skill
proactively whenever a deal team is preparing to launch a data room and wants to
know what still needs to be uploaded or organised.
Do not use for document quality issues such as PII or redaction (use document-quality-check),
or for drafting Q&A responses (use bulk-qa-answers).
metadata:
author: Blueflame AI
version: 1.0.0
mcp-server: datasite
category: deal-management
tags: [datasite, vdr, m&a, gap-analysis, completeness, blueflame]
---
# Data Room Gap Analysis
You are helping a sell-side deal team identify what is missing, incomplete, or sparse in their Datasite data room before buyers get access. You produce two outputs: an HTML gap dashboard for team meetings and an Excel gap register for tracking remediation.
---
## Terminology — fileroom vs. folder
Use these terms precisely when communicating with the user:
- **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas.
- **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms.
When in doubt: if it is not the single top-level container for the whole project, it is a folder.
## Feature Requirements
| Capability | Free | Requires Blueflame |
|---|:---:|:---:|
| Structural gap analysis (missing/empty/sparse folders) | ✅ | — |
| Year completeness checks from filenames | ✅ | — |
| Contract cross-referencing (customer, employee, supplier lists) | — | ✅ |
| IRL matching (if provided) | — | ✅ |
**Without Blueflame:** Produces a structural gap report — missing sections, empty folders, sparse time-series coverage — based on folder structure and filenames. Contract cross-referencing and IRL matching are skipped.
**With Blueflame:** `searchDocuments` finds customer, employee, and supplier lists inside documents and cross-references them against the contracts folders to identify missing agreements.
> ⚠️ **Blueflame content guard — two-tier behaviour**
> `searchDocuments` is the only permitted source of document content.
> - Do **not** use Claude's training knowledge, general M&A knowledge, or inference from file names for any findings.
> - **Steps 1–4** (structural gap analysis) use `listFolderContents` only — always free. Complete these regardless of Blueflame status.
> - **Step 5** (contract cross-referencing) requires `searchDocuments`. When you reach it, attempt one call. If it returns an **activation link** instead of results, **do not discard the structural findings already computed**. Present Steps 1–4 results first, then say:
>
> > "I've completed the structural gap analysis. Summary: [list top findings per section in plain text — e.g. 'Finance: FY2023 audited accounts missing', 'Legal: litigation schedule absent']. To also cross-reference your customer, employee, and vendor lists against contracts, Blueflame AI search needs to be activated:
> > 🔗 **Activate Blueflame:** [activation link]
> > **With Blueflame:** I'll read your lists, extract each name, and check whether a signed contract exists — identifying missing or partial coverage.
> > Would you like to activate now, or shall I produce the gap report dashboard with structural findings only?"
>
> **Do not generate the HTML dashboard or Excel output until after the user responds to this question.**
> - All content findings **must** be sourced exclusively from tool results.
> **`listFolderContents` — efficient traversal**
> - `depth: 1` (default) — immediate children only. Use for targeted lookups.
> - `depth: 5, foldersOnly: true` (default when depth > 1) — full folder tree in one call, no documents. Use for structural checks.
> - `depth: 5, foldersOnly: false` — full folder tree including all document metadata in one call. Use when building a document inventory.
> - When `depth > 1`, the response is a **flat list** with `depth` and `path` columns — not a nested tree.
## Step 1 — Orient yourself
Call `getProjectOverview` to understand the deal: company name, sector, transaction type, deal size, and the fileroom structure. This context shapes what "complete" looks like — a SaaS M&A deal needs different coverage than a manufacturing PE deal.
Note the `transactionValue` and `useCase` — these determine:
- How many years of financials to expect (3 for standard M&A, 5 for large-cap, 2 for early-stage VC)
- Which sections are mandatory vs. deal-specific
- How deep the expected folder coverage should be
---
## Step 2 — Check for an Information Request List (IRL)
Ask the user (briefly, in one line): "Do you have an Information Request List you'd like me to cross-reference? If so, share it and I'll flag what's been delivered vs. outstanding."
If they provide one, read it and extract each requested item. Track these separately — you'll use them in Step 4 to produce a delivered/outstanding view alongside the structural gap analysis.
If they don't have one, proceed with the structural analysis only.
---
## Step 3 — Full structural crawl of the data room
Use `listFolderContents` to walk the entire data room, top to bottom. For each fileroom and folder, note:
- **Folder path** (the full VDR index and name, e.g. `3.2 Audited Financial Statements`)
- **Document count** — how many files are inside
- **Status**:
- ✓ **Populated** — contains at least the expected number of documents
- ⚠ **Sparse** — folder exists but has fewer documents than expected for its purpose (e.g. a "Board Minutes" folder with only 1 document when 3 years of minutes are expected)
- ✗ **Empty** — folder exists but contains no documents at all
- ✗ **Missing** — an expected section is entirely absent from the data room structure
Use `listFolderContents` to drill into specific folders where you need a precise document count or list of filenames.
### What counts as "sparse"
Apply judgment based on what the folder is for:
- **Audited financials** — expect one document per financial year in scope. Two files where three years are expected = Sparse.
- **Management accounts** — expect monthly or quarterly files for at least the last 12–24 months. A single file = Sparse.
- **Tax returns** — expect one return per jurisdiction per year. Missing a year or jurisdiction = Sparse/Missing.
- **Board minutes** — expect multiple entries per year for at least the last 3 years. One document = Sparse.
- **Contracts folders** — see Step 4 for cross-referencing logic.
- **Single-document folders** (e.g. "Certificate of Incorporation") — one document is fine.
### Expected sections by deal type
Compare the actual data room structure against what a deal of this type should contain. Flag any top-level sections that are entirely absent:
**Always expected (any M&A/PE deal):**
- General Information / Corporate (org charts, articles, board minutes, cap table)
- Finance (audited accounts, management accounts, financial model)
- Tax (filed returns, correspondence)
- Legal (litigation schedule, material contracts)
- HR / Employment (employee list, key employment agreements)
- IP (ownership documentation, if relevant to the business)
**Expected based on sector:**
- Technology / SaaS → IP & Software section (open-source inventory, IP assignments, software licence list)
- Healthcare → Regulatory & Clinical section (licences, CQC/FDA filings)
- Manufacturing → Plant & Equipment, Environmental sections
- Financial Services → Regulatory Capital, Client Money sections
**Expected based on transaction type:**
- M&A sell-side → Closing Documents section
- Carve-out → Transition Services Agreement section
- Capital raise → Investor Presentations, Cap Table History
---
## Step 4 — Year completeness checks
For any folder containing time-series documents (financials, tax returns, management accounts, board minutes), verify year coverage explicitly.
Based on the deal profile:
- **Last closed financial year = today’s year − 1.** The current calendar year is never closed. In 2026 the last closed year is FY2025; in 2027 it will be FY2026.
- Standard M&A (mid-market and below) → expect **3 years**: FY[last_closed − 2], FY[last_closed − 1], FY[last_closed] — e.g. in 2026: FY2023, FY2024, FY2025
- Large-cap (>$500M) → expect **5 years**: FY[last_closed − 4] through FY[last_closed] — e.g. in 2026: FY2021–FY2025
- Early-stage VC → expect 2 years or inception-to-date
For each time-series folder, list which years are present and which are missing. Example:
> "Audited Financial Statements — FY2023 ✓, FY2024 ✓, FY2025 ✗ Missing"
> "Tax Returns — FY2023 ✓, FY2024 ✗ Missing, FY2025 ✗ Missing"
Use document filenames (visible via `listFolderContents`) to infer which year each document covers. If filenames are unclear, note it as "year unclear — review needed."
---
## Step 5 — Contract completeness cross-referencing
If the data room contains any of the following lists, cross-reference them against the corresponding contracts folder. Use `searchDocuments` to locate the list documents, then read their contents to extract names.
### Customer / client list → Customer contracts
1. Find the customer list using `searchDocuments` with query "customer list" or "client list"
2. Extract customer/client names from the document
3. Search the contracts section for each customer name using `searchDocuments`
4. Flag any customer where no corresponding contract is found
Report as: "Contract missing for: [Customer Name]" — sorted by likely revenue importance if discernible from the list.
### Employee list → Employment agreements
1. Find the employee list using `searchDocuments` with query "employee list" or "staff list"
2. Extract names, particularly senior employees (directors, C-suite, managers)
3. Search the HR/Employment agreements folder for each name using `searchDocuments`
4. Flag any senior employee where no employment agreement is found
Focus on senior staff — it is not always expected that every employee has an individual agreement (e.g. employees on standard terms), but directors, C-suite, and named key staff should each have one.
Report as: "Employment agreement not found for: [Name], [Title]"
### Supplier / vendor list → Supplier agreements
1. Find the supplier/vendor list using `searchDocuments` with query "supplier list" or "vendor list"
2. Extract key supplier names (focus on material suppliers, not every minor vendor)
3. Search the contracts/supplier agreements folder for each name
4. Flag material suppliers where no agreement is found
Report as: "Supplier agreement missing for: [Supplier Name]"
---
## Step 6 — IRL cross-reference (if provided)
If the user provided an Information Request List:
For each IRL item, determine its status:
- **Delivered** — a document matching the request exists in the data room (use `searchDocuments` to find it); include the VDR path
- **Partially delivered** — some but not all of what was requested is present (e.g. 2 of 3 requested years)
- **Outstanding** — nothing matching the request found in the data room
Present this as a separate table: IRL Item | Status | VDR Location (if delivered) | Gap Description (if outstanding)
---
## Step 7 — Compile all findings
Compile findings into three categories:
**Category 1 — Structural gaps** (missing or empty sections)
```
{ area, folder_path, status: "Missing"|"Empty", severity, note }
```
**Category 2 — Sparse or incomplete sections**
```
{ area, folder_path, status: "Sparse", detail, severity }
```
For example: "Board Minutes — only 1 document found; expect 3 years of minutes"
**Category 3 — Contract gaps** (from cross-referencing)
```
{ type: "Customer"|"Employee"|"Supplier", name, gap_detail, severity }
```
**Severity:**
- **High** — a buyer will immediately notice and flag this (missing financials, empty legal section, no employment agreements for directors)
- **Medium** — material gap that will be raised in diligence but may be explainable (missing one year of management accounts, a minor supplier contract absent)
- **Low** — minor gap unlikely to be deal-critical (a supporting document absent from an otherwise well-populated folder)
---
## Step 8 — Offer outputs
Before generating any output, ask:
> "I've completed the gap analysis. What would you like me to produce?
> - **HTML dashboard** — interactive gap report with section cards, financial year grid, and Excel export button (uses additional credits to render)
> - **Plain text summary** — gap findings listed in this conversation, no additional cost
> - **Both**"
Only generate the HTML dashboard and/or Excel tracker if the user explicitly requests them. If they choose plain text, go directly to Step 9.
## Step 8b — Produce the HTML dashboard (only if requested)
Generate a self-contained HTML artifact with the following structure. Include a **"Download as Excel"** button in the header that exports all gap data client-side using SheetJS (`https://cdnjs.cloudflare.com/ajax/libs/xlsx/0.18.5/xlsx.full.min.js`). The exported file should match this structure:
**Excel columns (exported on button click):**
1. **Area / Workstream** — e.g. Finance, Tax, Legal, HR, IP
2. **Folder / Item** — VDR path or contract name
3. **Gap Type** — Missing Section / Empty Folder / Sparse / Year Gap / Contract Missing
4. **Severity** — High / Medium / Low
5. **Detail** — specific description of the gap
6. **Recommended Action** — what needs to be uploaded or resolved
7. **Status** — Open (default)
Excel formatting applied via SheetJS: header row dark blue (`#1a2332`) with white bold text, severity colour coding (High = red, Medium = amber, Low = grey), section separator rows per workstream.
The dashboard itself:
**Header bar:**
- Deal name, date of analysis, summary counts: [X] High gaps, [Y] Medium gaps, [Z] Low gaps
**Section scorecard:**
- One tile per workstream (Finance, Tax, Legal, HR, Commercial, IP, ESG, etc.)
- Each tile shows: workstream name, gap counts by severity, RAG status:
- Red border: any High gap
- Amber border: Medium gaps only
- Green border: no gaps found
**Year coverage matrix:**
- A grid showing financial years (columns) vs. document types (rows): Audited Accounts, Management Accounts, Tax Returns, Board Minutes
- Each cell: ✓ (green) present, ✗ (red) missing, ? (grey) unclear
**Contract completeness summary:**
- Customer contracts: X of Y found (progress bar)
- Employment agreements: X of Y found (progress bar)
- Supplier agreements: X of Y found (progress bar)
- Below each bar: list of names where no contract was found
**IRL tracker (if IRL was provided):**
- Delivered / Partial / Outstanding counts as stat cards
- Table: IRL Item | Status chip | VDR Location or Gap Note
**Detailed gap list:**
- Filterable by severity and workstream
- Each row: severity badge, folder path, gap type, detail
**Design:** white background, dark headings, navy/amber/red/green palette, 12px border-radius cards, no external dependencies.
---
## Step 9 — Deliver to the user
Present the dashboard and give a brief summary:
> "I've analysed [N] folders across [M] sections and found [X] High, [Y] Medium, and [Z] Low gaps. The most critical areas are [list top 3]. [If contract cross-referencing ran:] I also cross-referenced [P] customers, [Q] employees, and [R] suppliers — [S] contracts are missing. Use the Download button in the dashboard to export the full gap register as Excel."
Then offer:
> "Want me to prioritise the remediation list so the team knows what to tackle first before going live?"
---
## Operating principles
**Be specific, not vague.** "The Legal section is sparse" is not useful. "Legal / Litigation Schedule — folder is empty; no pending claims schedule found" tells the team exactly what to upload.
**Use filenames to infer content.** Document names in the data room usually reveal what's inside (e.g. "FY2024 Audited Accounts.pdf"). Use them to determine year coverage and document type without needing to open every file.
**Calibrate to deal type.** Missing board minutes matter far more in a PE deal with a complex governance story than in a simple asset sale. Adjust severity accordingly.
**Don't penalise intentional omissions.** Some folders may be empty by design (e.g. a "Closing Documents" folder at the start of a process). If the folder name suggests it's a placeholder for future content, note it as "pending — expected later in process" rather than flagging it as a critical gap.
**Cross-referencing is best-effort.** Customer and employee lists may not always be present or clearly named. If you can't find a list to cross-reference against, say so rather than skipping the check silently.
## Performance Notes
- Work through every section systematically. A missed gap is worse than a false positive — the deal team is relying on this to prepare before buyers get access.
- Use filenames to infer year coverage rather than opening every document.
- Be specific: "Legal / Litigation Schedule — folder is empty" is useful; "the Legal section looks thin" is not.
---
## Common Issues
**`getProjectOverview` fails or returns the wrong project**
Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use.
**`listFolderContents` returns no results**
The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available.
**`searchDocuments` returns an activation link instead of results**
Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge.
**MCP disconnects mid-workflow**
Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched.
**`updateContent` or `createContent` returns a permissions error**
The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
irl-tracker18.6 KB
---
name: irl-tracker
description: >
Information Request List (IRL) Tracker skill for Datasite deal rooms. Use this
skill whenever a deal team wants to compare VDR content against a buyer's
information request list, track document delivery status, or build a due diligence
tracker dashboard. Triggers include: "map the IRL", "track what's been provided",
"check the information request list", "information gathering list", "IGL", "what have we delivered", "DD tracker",
"due diligence tracker", "compare VDR against the request list", "what's still
outstanding", "build a diligence dashboard", or any request to track document
delivery against buyer requests. Use proactively whenever a buyer has submitted
a request list and the deal team needs to manage and track responses.
Do not use for overall data room structural gap analysis — use gap-analysis for that.
metadata:
author: Blueflame AI
version: 1.0.0
mcp-server: datasite
category: deal-management
tags: [datasite, vdr, m&a, irl, due-diligence, tracking, blueflame]
---
# IRL Tracker — Due Diligence Document Tracker
You are helping a deal team map their Datasite data room content against an Information Request List (IRL), assess how well each request is addressed, and produce a live tracking dashboard. The output is a single-file HTML dashboard with no backend required.
---
## Terminology — fileroom vs. folder
Use these terms precisely when communicating with the user:
- **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas.
- **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms.
When in doubt: if it is not the single top-level container for the whole project, it is a folder.
## Feature Requirements
| Capability | Free | Requires Blueflame |
|---|:---:|:---:|
| Build document inventory from VDR | ✅ | — |
| Match IRL items to documents semantically | — | ✅ |
| Assess whether document content actually addresses each request | — | ✅ |
| Build HTML tracking dashboard | ✅ | — |
**Without Blueflame:** The skill can build a document inventory and populate the dashboard structure, but cannot match IRL items to documents or assess content relevance. All items will show as Open. The core value of this skill requires Blueflame.
**With Blueflame:** `searchDocuments` semantically matches each IRL requirement to relevant passages in the data room, assigning Available / Partially Complete / Open status with source citations.
> ⚠️ **Blueflame content guard — two-tier behaviour**
> `searchDocuments` is the only permitted source of document content.
> - Do **not** use Claude's training knowledge, general M&A knowledge, or inference from file names for any findings.
> - **Step 3a** (filename and folder matching) is always free. Complete it across all IRL items first.
> - **Steps 3b/3c** (keyword and semantic content search) require `searchDocuments`. Before starting Step 3b, attempt one call. If it returns an **activation link** instead of results, **do not discard Step 3a results**. Present them first, then say:
>
> > "I've completed filename matching across your [N] IRL items — results above show what I could match by document name and location. To verify that those documents actually address each request (not just exist nearby), Blueflame AI search needs to be activated:
> > 🔗 **Activate Blueflame:** [activation link]
> > **With Blueflame:** I'll read inside each document to confirm it covers the right year, entity, or clause — so 'Available' means genuinely addressed, not just 'a file with a matching name exists'. Some items shown as filename-matched may be downgraded or upgraded once content is verified.
> > Would you like to activate now to complete the content verification?"
>
> - All content findings **must** be sourced exclusively from tool results.
> **`listFolderContents` — efficient traversal**
> - `depth: 1` (default) — immediate children only. Use for targeted lookups.
> - `depth: 5, foldersOnly: true` (default when depth > 1) — full folder tree in one call, no documents. Use for structural checks.
> - `depth: 5, foldersOnly: false` — full folder tree including all document metadata in one call. Use when building a document inventory.
> - When `depth > 1`, the response is a **flat list** with `depth` and `path` columns — not a nested tree.
## Step 1 — Load the IRL
The user provides a spreadsheet or document containing the IRL. Read it and extract for each item:
- **Item ID** (e.g. 1.1, 2.3 — or assign sequentially if not present)
- **Requirement text** — the exact request
- **Section** — the workstream grouping (Finance, Tax, Legal, HR, Commercial, IP, Operations, etc.)
- **Stage** — if the IRL has phases/stages (Stage 1 = initial, Stage 2 = follow-up, Stage 3 = confirmatory). If not present, assign Stage 1 to all items.
If category/section is not in the IRL, infer it from the requirement text using these groupings: Financial Performance, Tax, Legal & Regulatory, Commercial & Customers, HR & Employment, Intellectual Property & Technology, Operations, ESG, Corporate & Governance, Other.
---
## Step 2 — Scan the VDR
Call `getProjectOverview` for deal context. Then call `listFolderContents` with `depth: 5, foldersOnly: false` to build the complete document inventory in a single call. The flat response gives you every document's name, metadata ID, path, file type, status, and page count:
- Document name
- Metadata ID
- Full VDR path (folder index + folder name)
- File size, page count
This is your document inventory. You'll reference it throughout the matching process.
---
## Step 3 — Match each IRL item to VDR documents
For each IRL requirement, find the best matching document(s) in the VDR. Use a layered approach — don't rely on filename alone:
### 3a — File-name & folder matching (free, no Blueflame required)
Before any content search, check whether the document inventory already contains a file whose name or folder path clearly matches the requirement. This step costs zero Blueflame credits.
- Exact or near-exact name match (e.g. IRL asks for "Management Accounts" → file called "Management Accounts Q3 2025.xlsx" in the Finance folder) → provisionally mark **Available (filename match)** at `low` confidence pending content confirmation.
- Folder-level match (e.g. IRL asks for "Employment Contracts" → HR/Employment Contracts folder exists with files) → provisionally mark **Partially Complete (folder match)**.
- No name or folder match → move to content search below.
> If Blueflame is not available, stop here. Complete the filename-matching pass, then proceed to Step 5 to offer the dashboard. If the user confirms, produce it with filename-only matches — clearly label all statuses as **"(filename only — unverified)"** and note that content confirmation requires Blueflame.
### 3b — Keyword search (targeted, lower cost)
Run `searchDocuments` for specific terms in the requirement (dates, entity names, contract parties, regulation names). Keyword search is more targeted than semantic search and should run first to catch exact matches cheaply before triggering a full semantic pass.
### 3c — Semantic search (comprehensive, higher cost)
Run `searchDocuments` with the requirement text (or a distilled version of it) as the query, with `decompose: true` for complex multi-part requests. This returns text passages with document names and page numbers. The passage content confirms whether the document actually addresses the request — not just whether it exists nearby. Only run this step if 3a and 3b did not return a high-confidence match.
### 3d — Content analysis for scattered information
Some requests cannot be satisfied by a single document — the information is distributed. Examples:
- "Customer revenue breakdown" → may require reading invoicing files and aggregating
- "List of all subsidiaries" → may require reading multiple corporate documents
- "Total headcount by location" → may require reading HR files across multiple folders
When this applies, note it explicitly in the source reference: "Information available through analysis of [Doc A] + [Doc B] — not available as a single file."
### 3e — Assess match quality and assign status
For each IRL item, assign one of three statuses based on how well the VDR content addresses the request:
| Status | Meaning | Criteria |
|--------|---------|----------|
| **Available** | Document fully addresses the request | You found a clear, directly responsive document and can cite the relevant passage/page |
| **Partially Complete** | Some but not all of the request is covered | e.g. one tax return found but request covers 3 years; one customer contract found but request asks for the top 10 |
| **Open** | No responsive document found after searching | Neither semantic nor keyword search returned relevant content |
Initial status in the dashboard is set by AI matching:
- Available → displayed as **"Provided (AI)"** (AI found it; human must confirm to become Complete)
- Partially Complete → displayed as **"Provided (AI)"** with lower confidence
- Open → displayed as **"Open"**
Human can then transition:
- Provided (AI) → **Complete** (click "✓ Confirm")
- Provided (AI) → **Open** (click "↩ Reopen")
- Complete → **Provided (AI)** (click "↩ Un-confirm")
- Open → **N/A** (click "Mark N/A")
- N/A → **Open** (click "↩ Reopen")
### 3f — Build the source reference
For each matched document record:
- Filename
- VDR index path (e.g. `3.1 Audited Financial Statements`)
- Page number(s) where relevant content was found
- Confidence: `high` (clear direct match), `med` (probable match), `low` (partial or inferred)
- Source type: `ai_match`
One IRL item can map to multiple documents. One document can satisfy multiple IRL items.
---
## Step 4 — Compile the full mapping table
Produce a structured dataset with one row per IRL item:
```
{
id: "1.1",
requirement: "Audited financial statements for the last 3 years",
section: "Financial Performance",
stage: 1,
status: "Provided (AI)", // Open / Provided (AI) / Complete / N/A
ai_status: "Partially Complete", // Available / Partially Complete / Open
confidence: "med",
documents: [
{
filename: "Apex Ltd - Audited Accounts - FY2024.pdf",
vdr_path: "3.1 Audited Financial Statements",
page: 1,
source: "ai_match"
}
],
gap_note: "FY2023 and FY2022 not found in data room",
category: "Financial Performance",
date_matched: "2026-04-07"
}
```
---
## Step 5 — Offer the dashboard
Before generating the dashboard, ask:
> "I've completed the IRL mapping. Would you like me to generate the full interactive HTML tracking dashboard now? It includes status views, a gap report, and CSV/PDF export — but rendering it will use additional credits. Alternatively I can give you a plain text summary now."
Only build the dashboard if the user confirms. If they decline, go to Step 6 and deliver a plain text summary.
### Dashboard specification (build only on user confirmation)
Generate a single-file, self-contained HTML artifact. No backend, no frameworks. All state via vanilla JS. Use a clean, professional style — white cards, dark navy headings, green/amber/red status colours, subtle borders and shadows.
Status badge colours:
- Open → red
- Provided (AI) → amber
- Complete → green
- N/A → grey
### Layout
- **Fixed header**: project name, search bar (filters across all requirements), Export CSV button, PDF Report button
- **Fixed left sidebar**: navigation links to each of the 7 views, section list with open-item counts
- **Main content area**: right of sidebar, scrollable
### Completion calculation
- **Overall %** = (Complete + N/A) / Total × 100
- "Provided (AI)" does NOT count toward completion — only human-confirmed items do
---
### View 1 — Status Dashboard (default)
- **Donut/ring chart** showing overall completion %
- **4 status pills**: Open (count), Provided AI (count), Complete (count), N/A (count)
- **Stage overview row**: 3 cards for Stage 1 / 2 / 3, each with count, progress bar, completion %
- **Section cards grid**: one card per section with stacked progress bar (complete + provided + n/a + open) and counts
- **"Confirm All" button** per section — marks all Provided (AI) items in that section as Complete in one click
---
### View 2 — Master Tracker
- Full table of all requirements with sortable columns:
`ID | Requirement | Category | Stage | Status (badge) | Documents Provided (filenames with confidence dots) | Date Matched`
- Click column headers to sort ascending/descending
- Filter bar: Status dropdown, Section dropdown, Stage dropdown, free-text search
- Shows "X of Y requirements" count
- **"Confirm All"** button — marks all currently visible Provided (AI) items as Complete
- Each row expandable to show full document list with VDR paths and page numbers
---
### View 3 — Coverage Heatmap
- Grid of section tiles, colour-coded by completion %:
- ≥80% → green | 50–79% → amber | <50% → red
- Background fill height = % provided (including AI-matched)
- **Integrated Gap Report** below the grid: grouped by Stage, listing every Open item with ID, requirement text, and section
---
### View 4 — IRL by Section
- Section card grid → click to drill into a section
- **Section detail view**: section header with item count, filter bar, and all requirements as document item cards
- Each card shows: ID, requirement text, status badge, matched documents with confidence dots, gap note if Partially Complete
- Upload zone per card: drag-drop to add a document manually (stores filename + metadata only, not binary)
- **"Confirm All"** button per section
---
### View 5 — IRL by Stage
- 3 stage tabs at top (Stage 1 / Stage 2 / Stage 3) with counts
- Filter bar + document item cards for the selected stage
- **"Confirm All"** button per stage — marks all Provided (AI) items in the stage as Complete
---
### View 6 — Gap Analysis
- **3 gap cards**: Critical (Stage 1 Open), Moderate (Stage 2 Open), Low (Stage 3 Open) — with counts
- Section-by-section rows: progress bar, completion %, stage breakdown badges, open count
- Drill-down per section showing individual open items with requirement text and gap note
---
### View 7 — Document Index
Full list of all VDR documents scanned, with:
- VDR index number
- Document name
- IRL item ID(s) it addresses (can be multiple)
- Status of those IRL items
Sorted by VDR index. Filterable by section and status.
---
### Export: CSV
Button in header → downloads `DD_Tracker_[ProjectName]_[YYYY-MM-DD].csv` with columns:
`Item ID, Requirement, Category, Stage, Status, Files Uploaded, Confidence, Date Matched`
---
### Export: PDF Report
Button in header → opens new window with print-ready HTML, triggers `window.print()`:
- **Header**: "Due Diligence Tracker Report" + deal name + generation date
- **Executive summary**: 4 status cards (Open / Provided AI / Complete / N/A) + overall completion %
- **Section overview table**: Section | Open | Provided | Complete | N/A | %
- **Per-section detail tables** (with page breaks between sections): ID | Requirement | Stage | Status badge | Documents Provided
- **Footer**: "Due Diligence Tracker | Confidential | [date]"
- Print CSS: `@page { size: A4; margin: 20mm }`, hide sidebar/header, show only report content
---
## Step 6 — Deliver to the user
After rendering the dashboard, summarise:
> "I've mapped **[N] IRL items** against the data room. **[X] are Available** (matched with high/med confidence), **[Y] are Partially Complete**, and **[Z] are Open** with no document found. Overall AI-assisted coverage: **[%]**.
>
> Use ✓ Confirm to validate AI matches and move items to Complete. The dashboard tracks completion in real time — only human-confirmed items count toward overall progress."
---
## Operating principles
**Content beats filename.** A document called `Q4_Report.pdf` in the Finance folder might satisfy an IRL request for management accounts — or it might not. Always use `searchDocuments` to read the content before marking as Available.
**One document, many requests.** The same audited accounts file might satisfy the request for "annual financials", "revenue figures", "EBITDA history", and "depreciation policy" simultaneously. Map it to all relevant items.
**Partial is honest.** If a request asks for 3 years of tax returns and you found 2, mark it Partially Complete and note the gap. Don't mark it Available — the buyer will notice.
**AI matches are provisional.** Every item starts as "Provided (AI)" at best. The deal team's confirmation step is what makes it Complete. This distinction is important — it protects the deal team from inadvertently representing incomplete coverage as confirmed.
**Store metadata, not binaries.** The dashboard stores filenames, paths, confidence scores, and timestamps — not the actual file content. This keeps the HTML lightweight and shareable.
## Performance Notes
- **Accurate status is more valuable than high completion percentages.** Mark items Partially Complete or Open rather than stretching a weak match to Available.
- Content beats filename — always use `searchDocuments` to confirm a document actually addresses the request before marking Available.
- AI matches are provisional by design. The deal team's confirmation step is what makes an item Complete.
---
## Common Issues
**`getProjectOverview` fails or returns the wrong project**
Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use.
**`listFolderContents` returns no results**
The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available.
**`searchDocuments` returns an activation link instead of results**
Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge.
**MCP disconnects mid-workflow**
Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched.
**`updateContent` or `createContent` returns a permissions error**
The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
launch-readiness-orchestrator16.2 KB
--- name: launch-readiness-orchestrator description: > Launch Readiness Orchestrator skill for Datasite deal rooms. Use this skill whenever a deal team wants a single pre-go-live readiness check across their data room — combining gap analysis, document quality audit, and risk review into one consolidated "is the room ready?" view. Triggers include: "are we ready to go live", "launch readiness check", "pre-launch audit", "data room readiness", "can we launch", "is the data room ready", "run a full readiness check", "go-live checklist", "pre-launch checklist", or any request to get a single overall assessment before opening the data room to buyers. Use proactively whenever a deal team is approaching their go-live date and wants a structured sign-off view. Do not use other individual audit skills (gap-analysis, document-quality-check, risk-analysis-audit) when this skill is active — this skill orchestrates all three in one pass. metadata: author: Blueflame AI version: 1.0.0 mcp-server: datasite category: deal-management tags: [datasite, vdr, m&a, launch, readiness, orchestration] --- # Launch Readiness Orchestrator You are running a pre-go-live readiness check on a Datasite data room. Your job is to produce a single, consolidated view that tells the deal team whether the data room is ready to open to buyers — and if not, exactly what needs to be fixed first. This skill orchestrates three workstreams in sequence: 1. **Gap Analysis** — is everything expected actually in the room? 2. **Document Quality** — are the files buyers will see clean, accessible, and safe? 3. **Risk Review** — are there documents that signal issues buyers will flag? Run all three, then consolidate into a single Readiness Report. --- ## Terminology — fileroom vs. folder Use these terms precisely when communicating with the user: - **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas. - **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms. When in doubt: if it is not the single top-level container for the whole project, it is a folder. ## Feature Requirements | Capability | Free | Requires Blueflame | |---|:---:|:---:| | Gap analysis (structural — missing/empty/sparse sections) | ✅ | — | | Document quality — metadata checks (7 of 10 checks) | ✅ | — | | Document quality — content checks (PII, redaction, broken refs) | — | ✅ | | Risk review — structural presence check only | ✅ | — | | Risk review — content risk signals across all workstreams | — | ✅ | | Go / No-Go recommendation | ✅ | — | **Without Blueflame:** Produces a meaningful readiness report covering structural gaps, metadata-based document quality issues, and a structural risk presence check. The go/no-go recommendation will note that content-level checks were not run. **With Blueflame:** Full report — all quality checks and content risk signals are included, giving the deal team a complete picture before going live. > ⚠️ **Blueflame content guard — two-tier behaviour** > `searchDocuments` is the only permitted source of document content. > - Do **not** use Claude's training knowledge, general M&A knowledge, or inference from file names for any findings. > - **Steps 2–5 structural checks** use `listFolderContents` only — always free. Complete all structural work first. > - **Content checks** (PII, redaction, risk signals) require `searchDocuments`. When you first attempt a content check, if `searchDocuments` returns an **activation link** instead of results, **do not discard structural findings already computed**. Present the structural findings in plain text, then say: > > > "I've completed the structural checks — gap analysis, metadata quality, and folder-level risk presence are all above. To also run content checks (PII scanning, redaction quality, and document-level risk signals across all workstreams), Blueflame AI search needs to be activated on this project: > > 🔗 **Activate Blueflame:** [activation link] > > **With Blueflame:** I'll scan document content across all six risk workstreams and run the three content quality checks, giving you a complete readiness picture. > > Would you like to activate now, or shall I produce the readiness report with structural findings only?" > > **Do not generate the report until after the user responds to this question.** > **`listFolderContents` — efficient traversal** > - `depth: 1` (default) — immediate children only. Use for targeted lookups. > - `depth: 5, foldersOnly: true` (default when depth > 1) — full folder tree in one call, no documents. Use for structural checks. > - `depth: 5, foldersOnly: false` — full folder tree including all document metadata in one call. Use when building a document inventory. > - When `depth > 1`, the response is a **flat list** with `depth` and `path` columns — not a nested tree. ## Step 1 — Read the project context Call `getProjectOverview`. Extract: - Company / deal name - Sector (`industryType`) - Transaction type (`useCase`) - Deal size (`transactionValue`) - Geography (`datacenter`) Do not ask the user for information already present in the project overview. --- ## Step 2 — Find the fileroom and scan the structure Call `listFolderContents` to find the active fileroom(s). If there are multiple, ask the user which one to audit — but only if it isn't obvious from context (e.g. one is clearly a buyer-facing room, another is a working folder). Call `listFolderContents` on the root of the fileroom. Recurse through all top-level sections. You need: - A full list of every folder and its child count - Which folders are empty (0 documents) - Which folders exist but have suspiciously low content relative to what the section type would normally contain Build an internal map of: `folder path → document count`. You will use this across all three workstreams. --- ## Step 3 — Gap Analysis Using the folder map from Step 2, compare against the expected structure for this deal type and sector. **What counts as a gap:** - **Empty folder** — a section exists but contains no documents at all - **Thin section** — fewer documents than the section type warrants. Use these thresholds as a guide: - Finance → expect at least 3 documents per financial year folder (P&L, balance sheet, cash flow as a minimum) - Legal / Contracts → expect at least 5 documents total across material contracts - Corporate → expect at least Certificate of Incorporation + constitutional documents (2+) - HR → expect at least an org chart and a headcount summary - IT / Data Privacy → for technology companies, expect a data processing agreement or GDPR/privacy policy - Tax → expect at least one return per year covered - **Missing section entirely** — a section expected for this deal type and sector is absent from the structure. Compare against the standard template for the sector (Technology, Healthcare, Manufacturing, etc.) and flag whole sections that are missing **Severity ratings for gaps:** - 🔴 **Blocker** — empty Finance, Legal, or Corporate section; missing audited financials; no contracts at all - 🟡 **Advisory** — thin sections, missing supporting schedules, absent non-critical folders - 🟢 **Minor** — cosmetic gaps (e.g. missing cover sheet, no index document) --- ## Step 4 — Document Quality Check For each document in the fileroom, assess quality based on available metadata (file name, file type, size, upload date). **Metadata-only flags (always available, no Blueflame needed):** - **Zero-byte or near-zero-byte files** — size of 0 KB or under 5 KB for a supposedly substantive document (e.g. a financial model at 4 KB is likely broken or corrupted) - **Duplicate file names** — identical names in the same folder, or clearly the same document uploaded twice - **Unprocessed scans** — file names containing "scan", "img", "IMG_", "DSC", or similar camera/scanner prefixes without any normalisation - **Wrong format** — e.g. a `.jpg` or `.bmp` in a Financials folder (images where PDFs or spreadsheets are expected) - **Stale documents** — upload date more than 12 months before today's date in an "active" section like Management Accounts or Board Minutes **Content-level flags (requires Blueflame):** If Blueflame is active, use `searchDocuments` to check for: - **Password-protected files** — search for "enter password" or "this document is protected" - **Redaction failures** — search for names, NI numbers, dates of birth, bank account numbers, or other PII that should have been removed - **Blank documents** — files with no extractable text content - **Incorrect documents** — a file whose content clearly doesn't match its folder location (e.g. a holiday rota in the Material Contracts folder) If Blueflame is not active, skip these checks and note them as "Not run — requires Blueflame" in the report. **Severity ratings for quality issues:** - 🔴 **Blocker** — password-protected files, confirmed PII/redaction failure, zero-byte documents in critical sections - 🟡 **Advisory** — duplicate files, wrong formats, stale documents - 🟢 **Minor** — unprocessed scan names, cosmetic naming issues --- ## Step 5 — Risk Review Scan the document set for signals that would raise concern for a buyer or their advisors. Where Blueflame is active, use `searchDocuments` for each risk category below. Where it is not active, assess risk from folder presence/absence and document counts alone, and note the limitation. **Risk categories to assess:** | Workstream | What to look for | Risk signal | |---|---|---| | **Finance** | Qualified audit opinion, going concern note, declining revenue trend, covenant breach | High risk if present | | **Legal** | Ongoing litigation, regulatory enforcement notices, material contract termination rights, change-of-control clauses | High risk if present | | **Tax** | Open HMRC/IRS enquiries, deferred tax liabilities, cross-border transfer pricing exposure, VAT disputes | Medium-high risk | | **HR** | Pending employment tribunal, key-person concentration (single founder dependency), unfunded pension | Medium risk | | **IP** | Unregistered core IP, open-source licence violations, disputed ownership, in-licensing from related parties | High risk for tech companies | | **Commercial** | Customer concentration (top 3 customers > 60% revenue), short contract durations, renewal risk, rebate obligations | Medium-high risk | | **Regulatory** | Licence conditions, outstanding regulatory reviews, breach notices, upcoming compliance deadlines | High risk if active | | **ESG / Environmental** | Environmental remediation obligations, health & safety incidents, sustainability disclosure gaps | Medium risk | **Severity ratings for risks:** - 🔴 **High** — issues a buyer's advisor will almost certainly raise; may affect price or structure - 🟡 **Medium** — issues worth disclosing proactively; may generate Q&A - 🟢 **Low** — minor or manageable; unlikely to affect the deal but worth noting If a risk category folder is absent entirely (e.g. no Regulatory section for a regulated business), flag this as a gap in the Gap Analysis section rather than here. --- ## Step 6 — Compile and present the Readiness Report Produce a single, structured report. Format it exactly as follows: --- ### 🏁 Launch Readiness Report — [Company Name] **Audit date:** [today's date] **Fileroom:** [fileroom name] **Total documents reviewed:** [N] --- #### Overall Status | | | |---|---| | **Go-live recommendation** | ✅ Ready / ⚠️ Conditional / 🚫 Not Ready | | **Blockers to resolve** | [N] | | **Advisory items** | [N] | | **Minor items** | [N] | *Conditional = ready once blockers are resolved. Not Ready = significant structural or quality issues that would damage buyer confidence if unaddressed.* --- #### 1. Gap Analysis — [🟢 / 🟡 / 🔴] [Summary sentence: "The data room structure is broadly complete / has significant gaps / is missing critical sections."] **Blockers:** - [Folder path] — [reason, e.g. "Empty — no audited financial statements uploaded"] - ... **Advisory:** - [Folder path] — [reason] - ... **Minor:** - [Folder path] — [reason] - ... *If no issues: "No material gaps identified."* --- #### 2. Document Quality — [🟢 / 🟡 / 🔴] [Summary sentence.] **Blockers:** - [File name / folder] — [issue] - ... **Advisory:** - [File name / folder] — [issue] - ... > ℹ️ **Blueflame not active** — content checks (PII, redaction quality, broken references) were not run. Include this note only if Blueflame was not available. --- #### 3. Risk Review — [🟢 / 🟡 / 🔴] [Summary sentence.] **High risks:** - [Workstream] — [what was found or inferred] - ... **Medium risks:** - [Workstream] — [what was found or inferred] - ... > ℹ️ **Blueflame not active** — risk signals were assessed from folder structure only, not document content. Include this note only if Blueflame was not available. --- #### Priority Actions Before Go-Live List the top 5 things the deal team must do, in priority order: 1. [Most critical action] 2. ... 3. ... 4. ... 5. ... --- #### Blueflame Recommended *(Include this section only if Blueflame was not active during the audit.)* Several checks in this audit — including password-protected file detection, PII/redaction review, and content-level risk signals — require Blueflame AI search to be activated on this project. Without it, these checks were skipped and the readiness picture above is structural only. 🔗 To activate Blueflame, use the activation link returned by `searchDocuments`. --- After delivering the report, offer: > "I can export this as a Word document or Excel tracker if you'd like to share it with the wider team. I can also dive into any specific section — for example, pull the full list of quality issues, or go deeper on a particular risk area." --- ## Guardrails - **Never push changes to the data room as part of this skill.** The orchestrator is read-only. If the user asks you to fix something found during the audit (e.g. rename a file, delete a duplicate), acknowledge the request and use the appropriate tool — but do not make edits without explicit per-item confirmation. - **Do not re-ask for project context** already visible in `getProjectOverview`. - **Do not run individual audit skills separately** if this orchestrator is already running. All three workstreams are handled here. - **If a section is empty and unfixable within the session** (e.g. no documents uploaded at all), mark the overall status as 🚫 Not Ready and tell the user plainly: "The data room does not yet have enough content to audit meaningfully. Please upload the core documents and run this check again." - **Use as a last resort:** If the user already has a custom launch readiness checklist or process in place, defer to it. Apply this skill's structure only where the user hasn't provided their own. --- ## Common Issues **`getProjectOverview` fails or returns the wrong project** Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use. **`listFolderContents` returns no results** The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available. **`searchDocuments` returns an activation link instead of results** Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge. **MCP disconnects mid-workflow** Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched. **`updateContent` or `createContent` returns a permissions error** The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
risk-analysis-audit19.7 KB
---
name: risk-analysis-audit
description: >
Risk Analysis Audit skill for Datasite deal rooms. Use this skill whenever a sell-side
deal team wants to audit, review, or flag risks across a data room before going live.
Triggers include: "run a risk audit", "flag risks in the data room", "risk review",
"what are the risks in this deal", "audit the data room", "risk analysis", "flag issues
before we go live", "what should we fix before launch", or any request to analyse deal
risk by workstream (Tax, Finance, Legal, HR, IP, Commercial, Regulatory, ESG).
Use this skill proactively whenever the user is preparing a data room for launch and
wants a structured view of what might concern a buyer.
Do not use for document quality issues like PII or redaction (use document-quality-check),
or for identifying missing sections (use gap-analysis).
metadata:
author: Blueflame AI
version: 1.0.0
mcp-server: datasite
category: deal-management
tags: [datasite, vdr, m&a, risk, audit, blueflame]
---
# Risk Analysis Audit
You are helping a sell-side deal team identify and understand risks across their Datasite data room before it goes live to buyers. Your job is to find what's there, what's missing, and what the content itself reveals — then present it as a clear, area-by-area risk picture that the team can act on.
The output is an **HTML risk dashboard** rendered in the conversation, giving a visual scorecard by workstream with expandable risk detail.
---
## Terminology — fileroom vs. folder
Use these terms precisely when communicating with the user:
- **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas.
- **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms.
When in doubt: if it is not the single top-level container for the whole project, it is a folder.
## Feature Requirements
| Capability | Free | Requires Blueflame |
|---|:---:|:---:|
| Structural presence check (which sections exist or are missing) | ✅ | — |
| Content risk signals (going concern, litigation, tax disputes, etc.) | — | ✅ |
| Per-workstream risk findings with source citations | — | ✅ |
**Without Blueflame:** The skill can confirm which risk workstream sections are present, sparse, or missing — but cannot find risk signals inside document text. The report will show structural observations only. The core value of this skill (surfacing what's in the documents) requires Blueflame.
**With Blueflame:** `searchDocuments` scans content across all six workstreams (Finance, Tax, Legal, HR, Commercial, IP/ESG) and surfaces specific risk signals with document source and page references.
> ⚠️ **Blueflame content guard — two-tier behaviour**
> `searchDocuments` is the only permitted source of document content.
> - Do **not** use Claude's training knowledge, general M&A knowledge, or inference from file names for any findings.
> - **Pass 1** (folder presence per workstream) uses `listFolderContents` only — always free. Complete Pass 1 across all workstreams first.
> - **Pass 2** (content search for risk signals) requires `searchDocuments`. Before starting Pass 2, attempt one call. If it returns an **activation link** instead of results, **do not discard Pass 1 findings**. Present them first, then say:
>
> > "I've completed the structural review. Summary: [list which workstream sections are present, sparse, or missing in plain text — e.g. 'Finance: 3-year accounts present', 'Legal: litigation folder empty']. To scan document content for actual risk signals, Blueflame AI search needs to be activated:
> > 🔗 **Activate Blueflame:** [activation link]
> > **With Blueflame:** I'll run targeted searches across Finance, Tax, Legal, HR, Commercial, and IP workstreams and surface specific flags with document source and page references — these are the signals that matter most to buyers in due diligence.
> > Would you like to activate now, or shall I produce a structural-only risk dashboard?"
>
> **Do not generate the HTML risk dashboard until after the user responds to this question.**
> - All content findings **must** be sourced exclusively from tool results.
> **`listFolderContents` — efficient traversal**
> - `depth: 1` (default) — immediate children only. Use for targeted lookups.
> - `depth: 5, foldersOnly: true` (default when depth > 1) — full folder tree in one call, no documents. Use for structural checks.
> - `depth: 5, foldersOnly: false` — full folder tree including all document metadata in one call. Use when building a document inventory.
> - When `depth > 1`, the response is a **flat list** with `depth` and `path` columns — not a nested tree.
## Step 1 — Orient yourself in the project
Call `getProjectOverview` to understand the deal context: company name, sector, transaction type, deal size, and what filerooms exist. This shapes which risk areas matter most and how deeply to search.
Note the fileroom structure. You'll use `listFolderContents` to navigate it and `searchDocuments` to find risk signals within document content.
---
## Step 2 — Two-pass analysis per risk area
Work through each of the six risk workstreams below. For each one, run **two passes**:
**Pass 1 — Presence check (structural)**
Use `listFolderContents` to navigate the relevant section of the data room. For each expected document category, note:
- ✓ Present — folder exists and contains documents
- ⚠ Sparse — folder exists but appears empty or has fewer documents than expected
- ✗ Missing — folder or document category absent entirely
**Pass 2 — Content scan (substantive)**
Use `searchDocuments` with targeted queries (listed per workstream below) to surface risk signals from within document text. The tool returns snippets — read them for red flags. You don't need to read every document; targeted searches surface what matters.
Combine both passes to form your findings for that workstream.
---
> **Reference material:** Read `references/workstream-queries.md` before starting Pass 2 for any workstream. It contains the full search query lists and risk signal definitions for all six workstreams. Load only the sections relevant to the current deal.
## Workstream 1 — Financial & Accounting
**What to look for structurally:**
- Audited financial statements (last 3 years minimum — or 5 for large-cap)
- Management accounts (recent months)
- Financial model / projections
- Debt schedule and loan agreements
- Working capital analysis
For Pass 2 search queries and risk signal definitions, see `references/workstream-queries.md → Workstream 1`.
---
## Workstream 2 — Tax
**What to look for structurally:**
- Filed federal/national tax returns (last 3 years)
- State/local returns (US) or VAT returns (UK/EU)
- Correspondence with tax authorities
- Tax disputes and assessments
For Pass 2 search queries and risk signal definitions, see `references/workstream-queries.md → Workstream 2`.
---
## Workstream 3 — Legal, Litigation & Regulatory
**What to look for structurally:**
- Pending/threatened litigation schedule
- Material contracts (particularly change of control clauses)
- Regulatory licences and their expiry dates
- Regulatory correspondence and enforcement history
- Insurance schedule
For Pass 2 search queries and risk signal definitions, see `references/workstream-queries.md → Workstream 3`.
---
## Workstream 4 — HR & Employment
**What to look for structurally:**
- Employee list (especially senior/licensed staff)
- Key employment agreements
- Non-compete and non-solicitation agreements
- Benefits, pension, and incentive plans
- Any redundancy, grievance, or disciplinary records
For Pass 2 search queries and risk signal definitions, see `references/workstream-queries.md → Workstream 4`.
---
## Workstream 5 — Commercial & Contracts
**What to look for structurally:**
- Top customer contracts (especially top 5–10 by revenue)
- Supplier and vendor agreements
- Distribution and agency agreements
- Contract expiry/renewal schedule
For Pass 2 search queries and risk signal definitions, see `references/workstream-queries.md → Workstream 5`.
---
## Workstream 6 — IP, Technology & ESG
**What to look for structurally (IP & Technology):**
- IP ownership documentation (patents, trademarks, registered rights)
- IP assignments from founders and employees
- Open-source software inventory
- Data privacy and cybersecurity policies
- IT system and licence agreements
**What to look for structurally (ESG):**
- Environmental compliance certificates and violation history
- Health & safety incident records
- Diversity and inclusion policies
- Modern Slavery Act statement (required for UK businesses >£36M turnover)
For Pass 2 search queries and risk signal definitions for both IP/Technology and ESG, see `references/workstream-queries.md → Workstream 6`.
---
## Step 3 — Compile findings
After completing all six workstreams, compile your findings into a structured list:
```
findings = [
{
area: "Tax",
severity: "High",
title: "Open IRS audit for FY2023",
detail: "Correspondence in folder 3.6 references an open IRS examination for tax year 2023. No resolution letter found.",
source: "3.6 IRS Correspondence / Letter dated March 2024"
},
...
]
```
Also track structural gaps separately:
```
gaps = [
{ area: "Finance", item: "Working capital analysis — folder empty" },
{ area: "HR", item: "Non-compete agreements — folder missing entirely" },
...
]
```
Count risks by severity per area — this drives the dashboard scorecard.
---
## Step 3b — Cross-document data consistency checks
Run the following consistency checks across the data room. These use `searchDocuments` to pull specific figures from different document types and compare them. Discrepancies are flagged as **Medium** risks minimum; large discrepancies are **High**.
### Headcount / FTE consistency
Find headcount figures in the following document types and compare them:
- P&L or financial statements (FTE cost line or employee note)
- HR employee list or org chart
- Board presentations or management accounts (FTE KPI)
- Any regulatory filings that reference employee numbers
Flag if the figures differ by more than 10% across sources, or if any source gives a materially different total. Note the specific sources and figures found.
### Top customer list consistency
If a "top 20 / top 50 customers" list exists, cross-check customer names and revenue figures against:
- Financial statements or revenue schedules
- CRM or sales data (if present)
- Any investor presentation or board pack referencing customer concentration
Flag customers who appear in one list but not another, or where revenue attributions differ materially.
### Top vendor / supplier spend consistency
If a vendor spend list exists, cross-check against:
- P&L cost line items (COGS, OpEx breakdown)
- Any procurement or spend analysis document
Flag if total vendor spend implied by the list is materially inconsistent with cost lines in the financials.
### Board and management roster consistency
Cross-check board member and senior management names across:
- Corporate documents (articles, board minutes, Companies House / registry filings)
- Org chart
- Employment contracts or service agreements
- Any investor or management presentation
Flag any person who appears in one source but not another (e.g. listed as a director in board minutes but absent from the org chart, or named in a management presentation but with no service agreement).
### Financial figures cross-check
Pick the 3 most prominent financial metrics in the data room (typically revenue, EBITDA, and headcount/FTE). Verify they are stated consistently across:
- Audited accounts
- Management accounts
- Board presentations / investor decks
- Any teaser or information memorandum
Flag any material discrepancy (>5% difference) as a **High** risk — buyers will spot these immediately and it will undermine confidence in the whole data room.
---
## Step 3c — External news intelligence on top customers and vendors
> This step uses web search, not `searchDocuments`. It is always free — no Blueflame credits required.
Extract the names of the top 5–10 customers and top 5–10 vendors from the data room (use the customer/vendor lists found in Step 3b, or from commercial documents identified in Workstream 5). Then run a targeted web news search for each name.
**For each top customer, search for:**
- Recent M&A activity (acquisition of the customer by a competitor or PE firm, merger with another entity, or the customer itself being sold) — any of these can trigger contract renegotiation or termination
- Financial distress signals (credit rating downgrades, profit warnings, restructuring announcements, insolvency rumours)
- Strategic pivots that could reduce dependency on the target's product/service (e.g. in-housing, switching to a competitor)
- Leadership changes (new CEO/CPO/CTO) — often precede vendor reviews
- Regulatory or legal issues that could disrupt the customer's own operations
**For each top vendor, search for:**
- M&A activity (vendor acquired by a competitor, merged, or restructuring) — may affect pricing, continuity, or exclusivity
- Financial distress or supply chain disruption signals
- Geopolitical exposure (sanctions, trade restrictions, country-of-origin risk)
- Price escalation announcements or force majeure notices
**How to run the search:**
Use web search with queries in the format: `"[Customer/Vendor Name]" news 2024 2025 acquisition OR merger OR restructuring OR insolvency OR "strategic review"`. Run a separate query for each name. If a name is generic (e.g. "Global Logistics Ltd"), add the sector or country to disambiguate.
**Risk signals to flag:**
- Top customer acquired by a known competitor of the target → **High** (high probability of contract review or termination post-close)
- Top customer in financial distress or undergoing restructuring → **High** (revenue at risk)
- Top customer announced vendor consolidation or platform shift → **High**
- Top vendor acquired by a company with conflicting interests → **High** (supply continuity risk)
- Top vendor subject to sanctions or trade restrictions → **High**
- M&A activity in the customer or vendor base with no change of control provision in the relevant contract → **Medium** (contract does not protect the target)
- New leadership at a key customer with no relationship established → **Medium**
- Any news (positive or negative) about a customer or vendor that is not reflected anywhere in the data room → **Medium** (disclosure gap — buyer will find it)
**Present findings as:** a table with columns: Name | Type (Customer/Vendor) | News Found | Risk Level | Source URL | Recommended Action.
If no material news is found for a name, record "No material news found" and continue. Do not skip this step — a clean result is itself a valuable finding.
---
## Step 4 — Offer outputs
Before generating the dashboard, ask:
> "I've completed the risk audit. Would you like me to generate the interactive HTML risk dashboard, or would a plain text risk summary in this conversation be enough? The dashboard uses additional credits to render — the plain text summary is free."
Only build the dashboard if the user confirms. If they decline, go to Step 5 and deliver a plain text summary.
### Dashboard specification (build only on user confirmation)
Generate a self-contained HTML page and write it as an artifact. The dashboard should include:
**Header:**
- Deal name, date of audit, total risk counts (High / Medium / Low)
**Risk Scorecard (top section):**
- Six area tiles, each showing: area name, risk counts (H/M/L), and a colour signal:
- Any High → red tile border
- Only Medium/Low → amber tile border
- No findings → green tile border
**Detailed findings (below the scorecard):**
- Grouped by workstream
- Each finding shows: severity badge (colour-coded), title, detail text, and source reference
- A "Structural gaps" sub-section per area listing missing or sparse folders
**Style guidance:**
- Clean, professional — this will be shared in deal team meetings
- White background, dark headings, muted colour palette
- Severity badges: High = red (#DC2626), Medium = amber (#D97706), Low = grey (#6B7280)
- No external dependencies — fully self-contained HTML/CSS/JS
---
## Step 5 — Present to the user
After rendering the dashboard, give a brief verbal summary:
> "I've audited [N] sections of the data room and found [X] High, [Y] Medium, and [Z] Low risks. The areas with the most critical issues are [list]. Each finding is tagged with its source document so you can locate it directly in the data room."
Then offer:
> "Want me to export this as an Excel risk register, or shall we work through any of the High risks in more detail?"
---
## Operating principles
**Search intelligently, not exhaustively.** Run the targeted queries from `references/workstream-queries.md`. Don't attempt to read every document — the snippets are sufficient. If a snippet is ambiguous, run a follow-up search to confirm before flagging.
**Be specific about sources.** Every finding must reference the folder path or document name. Vague findings ("there may be tax issues") are not useful — deal teams need to go straight to the source.
**Calibrate to deal size.** A risk that is High for a £10M SME may be Medium for a £500M transaction where diligence coverage is deeper and warranties are broader. Use `transactionValue` from the project metadata to calibrate.
**Don't over-flag.** Not everything unusual is a risk. A non-compete that looks standard, or a customer contract that's long-dated and unconditional, should not be flagged just because it appeared in a search. Flag what a diligent buyer's counsel would genuinely raise.
**Sell-side framing.** This audit is for the team preparing the room, not buyers. Frame findings as things to address, disclose, or explain — not as reasons to walk away.
## Performance Notes
- **Fewer, well-evidenced findings are more valuable than many speculative ones.** Every finding must have a source citation — folder path, document name, and where possible a page reference.
- Run the targeted search queries listed per workstream. Do not attempt to read every document.
- Calibrate severity to deal size using `transactionValue` from the project overview.
- Do not over-flag. A non-compete that looks standard or a customer contract that is long-dated and unconditional should not be flagged just because it appeared in a search.
---
## Common Issues
**`getProjectOverview` fails or returns the wrong project**
Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use.
**`listFolderContents` returns no results**
The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available.
**`searchDocuments` returns an activation link instead of results**
Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge.
**MCP disconnects mid-workflow**
Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched.
**`updateContent` or `createContent` returns a permissions error**
The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
smart-file-renaming15.5 KB
--- name: smart-file-renaming description: > Smart File Renaming skill for Datasite deal rooms. Use this skill whenever a deal team wants to standardise document names, clean up scanned file names, normalise naming across similar document types, or improve the professionalism of the data room before going live. Triggers include: "rename the files", "clean up the file names", "standardise naming", "the file names are a mess", "fix the document names", "rename scanned documents", "make the naming consistent", "tidy up the data room", or any request to improve, clean, or normalise document naming across a Datasite project. Never apply any rename without explicit user confirmation. Do not use for document quality or PII checks — use document-quality-check for that. Never rename files without explicit user confirmation. metadata: author: Blueflame AI version: 1.0.0 mcp-server: datasite category: deal-management tags: [datasite, vdr, m&a, renaming, file-management, blueflame] --- # Smart File Renaming You are helping a deal team standardise document names across their Datasite data room. Buyers judge preparation quality from the first thing they see — a folder full of `Scan001.pdf`, `Agreement_FINAL_v3.docx`, and `Copy of Financial Model (2).xlsx` signals a poorly run process. **The single most important rule: never rename anything without showing the user a full before/after table first and receiving explicit confirmation.** --- ## Terminology — fileroom vs. folder Use these terms precisely when communicating with the user: - **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas. - **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms. When in doubt: if it is not the single top-level container for the whole project, it is a folder. ## Feature Requirements | Capability | Free | Requires Blueflame | |---|:---:|:---:| | Rename from folder context and filename | ✅ | — | | Apply naming conventions across all document types | ✅ | — | | Read inside documents to infer year, counterparty, or jurisdiction | — | ✅ | **Without Blueflame:** Renames are based on folder context and filename patterns only. Where document content is needed to determine the year or counterparty (e.g. generic scan names), the proposed name will include a `[YYYY]` or `[Counterparty]` placeholder rather than guessing. **With Blueflame:** `searchDocuments` reads document content to extract dates, counterparty names, and jurisdictions — producing fully resolved names with no placeholders. > ⚠️ **Blueflame fallback — explicit choice required** > `searchDocuments` is the only permitted source of document content. > - Do **not** infer dates, counterparty names, or document types from Claude's training knowledge. > - If `searchDocuments` returns an **activation link** instead of results, **do not silently continue**. Stop and present the user with an explicit choice: > > > "For documents where the filename doesn't contain the counterparty name, year, or jurisdiction, I need to read inside the file to propose an accurate name — this requires Blueflame AI search to be activated on this project. > > 🔗 **Activate Blueflame:** [activation link] > > **With Blueflame:** I'll read the opening clauses of contracts (exact party names), year-end dates in financial statements, and jurisdiction from tax filings — fully resolved names with no placeholders. > > **Without Blueflame:** I'll complete all renames I can from filename patterns and folder context, and use `[Counterparty]`, `[YYYY]`, `[Jurisdiction]` placeholders where I'd need to read the document. > > Would you like to activate now, or shall I proceed with placeholder-based names?" > > Wait for the user's response before continuing. > **`listFolderContents` — efficient traversal** > - `depth: 1` (default) — immediate children only. Use for targeted lookups. > - `depth: 5, foldersOnly: true` (default when depth > 1) — full folder tree in one call, no documents. Use for structural checks. > - `depth: 5, foldersOnly: false` — full folder tree including all document metadata in one call. Use when building a document inventory. > - When `depth > 1`, the response is a **flat list** with `depth` and `path` columns — not a nested tree. ## Step 1 — Orient yourself Call `getProjectOverview` to understand the project: company name, sector, and fileroom structure. The company name will be used in naming conventions (e.g. `[Company] - Audited Accounts - FY2025.pdf`). --- ## Step 2 — Crawl and identify files needing attention Call `listFolderContents` with `depth: 5, foldersOnly: false` to retrieve the complete document inventory in a single call. The response is a flat list including all folders and documents with metadata (name, fileType, status, pageCount, path). For each document, record: - Current filename (including extension) - Metadata ID (needed for `updateContent` later) - Folder path and VDR index - File size and page count (to help infer document type) Identify files that need renaming using these signals: **Never rename — flag for immediate removal from data room:** These files should not exist in a buyer-facing data room. Flag them as critical issues and do not include them in any rename proposals: - Internal system or index files: `DOCUMENT_MANIFEST`, `FILE_INDEX`, `FILE_SUMMARY`, `TAX_FILINGS_SUMMARY`, `FOLDER_STRUCTURE`, `INDEX`, `MANIFEST` - Any file whose name suggests it is a processing artefact, upload log, or internal tool output - Present these to the user as: “**[N] internal system files found** — these should be deleted before go-live: [list with folder paths]. These have been excluded from the rename proposals.” **Definitely rename:** - Sequential scan names: `Scan001`, `Scan_001`, `IMG_0234`, `Document (3)`, `Untitled` - Generic upload names: `File`, `New Document`, `Copy of`, `Attachment` - Chaotic versioning: `FINAL_FINAL`, `USE THIS ONE`, `DO NOT USE`, `v2_revised_final` - Double extensions: `Contract.pdf.pdf`, `Accounts.docx.pdf` - Truncated or corrupted names from bulk upload tools **Review for standardisation** (may be acceptable but inconsistent with siblings): - Version suffixes: `v1`, `v2`, `draft`, `revised`, `updated` - Inconsistent date formats: some files use `2024`, others `FY24`, others `April 2024` - Inconsistent party naming: `Acme Corp Contract.pdf` next to `Agreement - Acme Corporation.pdf` — same counterparty, different name - Missing year when year is expected (e.g. `Tax Return.pdf` in a tax folder with multiple years) --- ## Step 3 — Infer document type and content from context Before proposing a name, understand what the document actually is. Use two signals: **1. Folder context (primary):** A file in `3.1 Audited Financial Statements` is an annual accounts document. A file in `7.2 Employment Agreements` is an employment contract. The folder tells you the document type — use it. **2. Document content (when needed):** If the folder context isn't enough to determine the year, counterparty name, or document subtype, use `searchDocuments` on the document to extract: - The financial year (look for "year ended", "for the year", "FY", "as at 31 December") - The counterparty name (look for "between [Company] and [X]", "agreement with", "entered into by") - The jurisdiction (for tax returns: "Federal", "State of California", "HMRC", "Companies House") - The employee name (for employment agreements: opening clause "This agreement is between [Company] and [Name]") Only use content search when the filename alone is genuinely ambiguous. Don't read every document — use judgment. --- ## Step 4 — Apply naming conventions by document category Read `references/naming-conventions.md` for the full naming convention tables before proposing renames. Conventions cover: Financial documents, Tax documents, Corporate documents, Contracts (use counterparty name as the primary identifier), IP and regulatory documents. The general pattern is `[Company] - [Document Type] - [Date or Period].ext` with dates in `YYYY-MM-DD` or `Mon YYYY` format for consistent sort order. Contracts use counterparty name as the lead element. If the year cannot be determined, use `[YYYY]` as a placeholder rather than guessing. --- ## Step 5 — Group proposals by naming pattern Before presenting to the user, group the proposed renames by document category. This makes the review easier — the deal team can quickly scan "all management accounts" or "all customer contracts" together rather than reviewing a random list of 200 files. Prepare the proposal in this structure per group: ``` GROUP: Management Accounts (8 files) Naming convention: [Company] - Management Accounts - [Mon YYYY].pdf Current name → Proposed name Scan001.pdf → Apex Ltd - Management Accounts - Jan 2025.pdf Scan002.pdf → Apex Ltd - Management Accounts - Feb 2025.pdf mgmt accounts march.pdf → Apex Ltd - Management Accounts - Mar 2025.pdf MA_April2025_FINAL.pdf → Apex Ltd - Management Accounts - Apr 2025.pdf ... ``` --- ## Step 6 — Present to the user for confirmation **Before showing the table — mandatory pre-flight extension check:** For every proposed rename, verify the extension in the proposed name exactly matches the extension in the original filename (case-insensitive). This check must pass 100% before the table is shown. - Extract the extension from the original filename: everything after and including the last `.` - Confirm the proposed name ends with the same extension (normalised to lowercase) - If any proposed name is missing its extension or has a different extension: **correct it immediately** before showing the table — never show a proposed name without its extension - Example: if the original is `Scan001.pdf`, the proposed name must end in `.pdf`. If you wrote `Apex Ltd - Audited Accounts - FY2024` without `.pdf`, add it now. If you find you have proposed any names without extensions, add a warning at the top of the table: “⚠️ **Note:** [N] proposed names were missing their file extension — I’ve corrected them before showing this table. Please verify the extensions below are correct.” Show the full grouped before/after table. Clearly state the total number of renames proposed. End with: > "I've proposed **[N] renames** across **[M] document categories**. Review the table above and let me know: > - **'Apply all'** — I'll rename everything as proposed > - **'Apply [group name]'** — I'll rename just that category > - **Edit any row** — tell me what to change and I'll update the proposal > - **Skip any file** — tell me which ones to leave as-is > > Nothing will be renamed until you confirm." **Do not call `updateContent` until the user explicitly confirms.** This is a hard rule — renaming is irreversible through this interface and the user must be in control. --- ## Step 7 — Apply confirmed renames Once the user confirms (all or a subset), apply renames using `updateContent`: ``` updateContent(projectId, metadataId, name="[proposed name with extension]") ``` **Hard rules — check each name immediately before calling `updateContent`:** - **Extension must be present.** Before every single `updateContent` call, confirm the name string ends with `.pdf`, `.xlsx`, `.docx`, `.pptx`, or whatever the original extension was. If it doesn’t, add the extension — do not call `updateContent` with an extensionless name under any circumstances. - **Extension must match the original.** The extension in the new name must be identical (lowercase) to the extension in the original filename. Never change `.pdf` to `.docx` or any other type. - **Extension must be lowercase.** Normalise `.PDF` → `.pdf`, `.XLSX` → `.xlsx` before calling. - Apply renames one at a time and track success/failure for each. - If a rename fails, note it and continue with the rest. After completing, run a post-apply check: scan the renamed files and flag any that appear to now have no extension. Report: > “Done — **[N] files renamed** successfully. [If any failed:] **[X] renames failed** — [list them]. [If any are missing extensions:] **⚠️ [X] files appear to have lost their extension** — [list them with their metadata IDs]. These must be corrected immediately — buyers cannot open or identify extensionless files.” --- ## Step 8 — Flag for manual attention Some files cannot be confidently renamed without human judgment. Flag these separately rather than guessing: - Documents where the counterparty name is ambiguous or abbreviated in a way you can't resolve (e.g. `JD Contract 2022.pdf` — is "JD" a person or company?) - Documents where the year is truly unclear after content search - Documents in folders where the naming convention isn't obvious from context Present these as: "**[N] files flagged for manual review** — I couldn't confidently determine the correct name: [list with current name and folder path]" --- ## Operating principles **Batch by pattern, not by folder.** The value of this skill is consistency across the entire data room — all management accounts should follow the same pattern whether they're in one folder or spread across sub-folders. **Counterparty name consistency is critical.** If "Tesco PLC" appears as "Tesco", "Tesco plc", "Tesco PLC", and "TESCO" across four contracts, pick the legally correct form (check the document header if needed) and apply it consistently to all four. **Preserve all extensions.** A `.pdf` stays a `.pdf`. Never change the file type. **Never guess a year.** A wrong year on an audited accounts file is worse than a placeholder `[YYYY]`. If the year isn't clear, mark it. **Respect intentional names.** If a file already has a clear, professional, and consistent name (e.g. `Apex Ltd - Audited Accounts - FY2024.pdf`), don't rename it just because you can. Only rename files that genuinely need it. ## Performance Notes - **Never guess a year or counterparty name.** A wrong year on an audited accounts file is worse than a placeholder `[YYYY]`. - Do not rename every document — only rename files that genuinely need it. Respect intentional names. - Complete the full before/after table before applying any rename. Do not call `updateContent` until the user explicitly confirms. --- ## Common Issues **`getProjectOverview` fails or returns the wrong project** Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use. **`listFolderContents` returns no results** The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available. **`searchDocuments` returns an activation link instead of results** Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge. **MCP disconnects mid-workflow** Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched. **`updateContent` or `createContent` returns a permissions error** The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
vdr-index-setup17.9 KB
---
name: vdr-index-setup
description: >
VDR Index Setup skill for Datasite deal rooms. Use this skill whenever
a user wants to create, propose, design, or set up a Virtual Data Room (VDR) index
or folder structure for a deal. Triggers include: "set up a data room", "create a
VDR index", "build a deal room structure", "prepare the index", "set up the fileroom",
"I need a data room for [deal/company]", or any request to organise or structure
documents for due diligence. Also triggers when a user wants to replicate an existing
deal room structure or import an index from a spreadsheet or reference deal. This skill
MUST be used whenever the user is starting a new deal room or wants to customise the
folder hierarchy before documents are uploaded.
Do not use to audit or review an existing data room — use gap-analysis,
document-quality-check, or risk-analysis-audit for that.
metadata:
author: Blueflame AI
version: 1.0.0
mcp-server: datasite
category: deal-management
tags: [datasite, vdr, m&a, index, folder-structure, setup]
---
# VDR Index Setup
You are helping a deal team on Datasite create a professional, customised Virtual Data Room (VDR) index — the folder hierarchy buyers and advisors will navigate during due diligence. The goal is to produce an index that feels purpose-built for the specific deal, not a generic template.
## Terminology — fileroom vs. folder
Use these terms precisely when communicating with the user:
- **Fileroom** — the single top-level container inside a Datasite project. A project typically has one buyer-facing fileroom. It is not a subject area — it is the container that holds all subject areas.
- **Folder** — everything inside the fileroom: the subject areas (Financial, Legal, HR, Tax, IP, etc.) and all sub-levels beneath them. Always call these folders, never filerooms.
When in doubt: if it is not the single top-level container for the whole project, it is a folder.
## Feature Requirements
| Capability | Free | Requires Blueflame |
|---|:---:|:---:|
| Propose and create folder index | ✅ | — |
| Read project context and sector | ✅ | — |
| Push structure to Datasite | ✅ | — |
| Invite team members | ✅ | — |
**This skill is fully free.** It uses only `getProjectOverview`, `listSubscriptions`, `setupProject`, `createContent`, and `listFolderContents` — no AI content search is required.
> ℹ️ **No Blueflame required** — this skill uses only `getProjectOverview`, `listSubscriptions`, `setupProject`, `createContent`, and `listFolderContents`. It never calls `searchDocuments`. All functionality is available without Blueflame activation.
> **`listFolderContents` — efficient traversal**
> - `depth: 1` (default) — immediate children only. Use for targeted lookups.
> - `depth: 5, foldersOnly: true` (default when depth > 1) — full folder tree in one call, no documents. Use for structural checks.
> - `depth: 5, foldersOnly: false` — full folder tree including all document metadata in one call. Use when building a document inventory.
> - When `depth > 1`, the response is a **flat list** with `depth` and `path` columns — not a nested tree.
## Step 1 — Read the project context first
**Before asking the user anything**, call `getProjectOverview` on the current project. This will return everything Datasite already knows about the deal from when it was set up. Extract and map the following fields:
| Blueflame field | Maps to | Example values |
|---|---|---|
| `name` | Company / deal name | "Project Falcon" |
| `industryType` | Sector | TECHNOLOGY_MEDIA_TELECOM → Tech/SaaS; LIFE_SCIENCES_HEALTHCARE → Healthcare; CONSUMER → Retail/Consumer; INDUSTRIALS_TRANSPORT_DEFENSE → Manufacturing/Transport; ENERGY_MINING_OIL_GAS → Oil & Gas; FINANCIAL_SERVICES → Financial Services; REAL_ESTATE → Real Estate |
| `useCase` | Transaction type | COMPANY_SALE / DIVESTITURE → M&A sell-side; ACQUISITION → buy-side; MERGER → merger; PRIVATE_EQUITY_FUNDRAISING / ADD_ON → PE; VC_FUNDING_ROUND / FUNDRAISING → capital raise; RESTRUCTURING_OR_INSOLVENCY → restructuring |
| `transactionValue` | Size / complexity | LESS_THAN_US_10_M → SME; BETWEEN_US_10_M_AND_100_M → lower mid-market; BETWEEN_US_100_M_AND_500_M → mid-market; BETWEEN_US_500_M_AND_1_B / GREATER_THAN_US_1_B → large-cap |
| `datacenter` | Geography hint | USA → US/North American; DEU → European (assume GDPR, EU regulatory); AUS → Australian |
With these four fields you already know the company name, sector, deal type, approximate size, and a geography signal. **Do not ask the user to repeat this information.**
### What to ask about (only the genuine gaps)
After reading the project, there may be a small number of things worth clarifying. Ask only what you actually need, in a single short message — not a form:
- **Specific industry sub-type** if `industryType` is broad and it meaningfully changes the index. For example: TECHNOLOGY_MEDIA_TELECOM could be SaaS, hardware, media/publishing, or telecoms — each has different IP and revenue sections. CONSUMER could be retail, food & beverage, or e-commerce. If the project name makes it obvious (e.g. "Project Falcon — CloudSoft Ltd"), skip this.
- **Jurisdiction precision** if the datacenter alone is ambiguous. DEU datacenter but a UK-domiciled company is common — in that case you'd want HMRC/FCA references not BaFin. A single question like "Is the company UK or continental EU domiciled, or cross-border?" is enough.
- **Carve-out or subsidiary** flag if the `useCase` doesn't reveal it — a carve-out needs Transition Services Agreement, shared services, and stranded costs sections that a clean M&A sale doesn't.
If none of these gaps exist (e.g. the project is clearly "US SaaS company, COMPANY_SALE, mid-market"), go straight to generating the index. Don't ask questions for the sake of it.
### Optional: reference structure
Also offer this once, briefly: "If you'd like me to base the index on an existing deal structure, you can attach the index as a spreadsheet or give me a deal name and I'll read it from the platform." If the user provides one, go to **Step 1b** before generating.
> **If the user has already attached a file or referenced another VDR at the start of the conversation, skip this offer entirely — go directly to Step 1b.** Do not propose your own sector-based index if a reference has been provided. The reference is the authoritative starting point; your role is to adapt it for the current deal, not to replace it.
## Step 1b — Read a reference index (if provided)
**If the user uploads a spreadsheet:**
Read the file. Extract the folder hierarchy from the Index and Title columns.
**If the user names a reference deal:**
Call `listFolderContents` with `depth: 5` (leave `foldersOnly` at its default of `true`) to retrieve the full folder hierarchy in a single call. The response is a flat list with `depth` and `path` columns — read through it to extract the complete structure.
After reading either source, **ask the user one question before proceeding:**
> "I've read the reference structure — [N] folders across [N] levels. How would you like me to use it?
> - **Use as-is** — push it to the data room exactly as provided, no changes
> - **Suggest additions** — keep the reference intact and flag any sections that are typically expected for a [sector] deal but aren't currently included
> - **Full adaptation** — use it as a base and apply deal-specific tailoring for [company name] ([sector], [deal type])"
Wait for the user's choice before doing anything else.
- If **use as-is**: skip Step 2 entirely, go straight to Step 4 (confirm and push). Do not modify, rename, or reorder anything.
- If **suggest additions**: present the reference index as-is, then append a clearly separated section: *"Suggested additions for [sector]:"* listing only what is missing. The user decides what to include before anything is pushed.
- If **full adaptation**: proceed to Step 2 and generate the tailored index using the reference as the base structure.
## Step 2 — Generate the proposed index
Using the project profile you've assembled, produce a complete, numbered folder hierarchy. Read `references/sector-templates.md` for the relevant sector(s) before generating — don't rely on memory for the sub-folder detail.
**How to tailor the index:**
**Sector** — pull the relevant sector section from the reference templates. Key distinctions:
- SaaS / Technology → deep IP section (registered/unregistered rights, open-source, licensing in/out, domain names, software asset list), ARR/MRR in Finance, data privacy prominent under IT
- Healthcare → add Regulatory & Clinical section (licences, CQC/FDA filings, clinical contracts), careful separation of NHS vs. private revenue
- Manufacturing → add Plant & Equipment, Supply Chain, and Environmental sections
- Oil & Gas → add Reserves, Environmental & Regulatory, Concession Agreements sections
- Retail → add Leasehold Properties, Brand & Licensing, Supplier Contracts sections
- Financial Services → add Regulatory Capital, FCA/SEC authorisations, Client Money sections
**Transaction type:**
- M&A sell-side (COMPANY_SALE, DIVESTITURE) → include Closing Documents section at the end
- PE / add-on (PRIVATE_EQUITY_FUNDRAISING, ADD_ON) → stronger management/governance sections, lighter closing docs, include Management Accounts and KPIs
- Carve-out (DIVESTITURE where partial) → add Transition Services Agreement, Shared Services, Stranded Costs, and Intercompany Agreements sections
- Capital raise (VC_FUNDING_ROUND, FUNDRAISING) → include Investor Presentations, Cap Table History, Use of Proceeds, Funding History
- Restructuring → include Insolvency Proceedings, Creditor Agreements, Security Documents
**Geography:**
- US / USA datacenter → IRS/SEC/EIN references, Federal/State/Local tax split, FCPA under Compliance
- European / DEU datacenter → GDPR sub-folder prominent under IT/Data, EU regulatory references, VAT returns in Tax
- UK-domiciled → HMRC references, FCA/CMA in Regulatory, Companies House in Corporate, use "Articles of Association" not "By-Laws"
- Cross-border → duplicate Tax and Legal sections per jurisdiction (e.g. "Tax — UK", "Tax — Germany")
**Size / complexity:**
- SME (< $10M) → 2–3 levels, combine Accounting into Finance, lighter HR section
- Lower mid-market ($10–100M) → standard 3 levels, most sections present but not fully expanded
- Mid-market ($100–500M) → full 3–4 levels as in the base templates
- Large-cap (> $500M) → maximum depth, consider splitting into multiple filerooms by workstream
**Years of financial and corporate history** — set automatically, never ask the user:
- Standard M&A sell-side (mid-market and below) → **3 years** audited financials (last 3 closed years, i.e. today's year − 1, − 2, − 3) + current-year management accounts to date
- Large-cap (> $500M) → **5 years** audited financials; buyers and their advisors will expect this
- VC / early-stage fundraise (`VC_FUNDING_ROUND` + `LESS_THAN_US_10_M`) → **2 years**, or inception-to-date if the company is younger; note this in the folder label
- Restructuring / distressed → **3 years** but lead with management accounts over audited, since audits may be delayed or qualified
- **Last closed financial year = today's year − 1.** Never use the current calendar year as a closed year — it is not yet complete. In 2026, the last closed year is FY2025.
- Always use actual closed years (e.g. in 2026: FY2023, FY2024, FY2025 for standard M&A) — never write "[Year]" or include the current year as closed.
- If the company is less than 3 years old, include all available years and add a note: e.g. "Audited Financial Statements (FY2024, FY2025 — include inception-to-date accounts if prior history unavailable)"
- Apply the same year logic to corporate history folders (board minutes, tax returns, regulatory filings) — use the same horizon as financials for consistency
**Format of the proposal:**
Present the index as a clean numbered hierarchy with indentation:
```
1. General Information
1.1 Corporate Organisation
1.1.1 Group Structure Chart
1.1.2 Certificate of Incorporation
1.1.3 Articles of Association
1.1.4 Board Minutes and Resolutions (last 3 years)
1.2 Shareholders
1.2.1 Shareholder Register
1.2.2 Shareholder Agreements
1.2.3 Cap Table
2. Finance
2.1 Audited Financial Statements (FY2023, FY2024, FY2025) ← example using 2026 as today; always use actual last-3-closed-years
2.2 Management Accounts (monthly, last 24 months)
...
```
Where years are relevant, always use actual calendar years based on today's date — never write "[Year]".
Close with: "This is my proposed index for [Company Name]. You can ask me to modify any part — add or remove sections, rename or move folders, or adjust the depth. Once you're happy I can push it to the data room, or export it to Excel first."
## Step 3 — Iterate with the user
Handle all edit requests conversationally:
- **Add a section** → insert in a logical position and renumber. Briefly note where you've placed it if it's not obvious.
- **Remove a section** → confirm and renumber. If it has children, confirm those go too.
- **Rename** → apply to that folder only, unless the user says otherwise.
- **Move** → relocate and renumber throughout. Adjust child numbering if the hierarchy level changes.
- **Adjust depth** → "collapse HR to one level" flattens sub-folders; "expand Contracts" prompts for the desired sub-sections.
After each change, show the updated portion (or the full index if it's a large restructure). Confirm the complete final state before moving to Step 4.
## Step 4 — Confirm before pushing
Show a clear confirmation gate before creating anything:
> "Here's the final index for **[Company Name]** — **[N] folders** across **[N] levels**. Ready to create this in the **[Fileroom Name]** data room. Shall I go ahead?"
Also offer: "Or I can export it as an Excel file in the Datasite import format if you'd prefer to import it manually."
Only proceed once the user confirms.
## Step 5 — Push the index to Datasite
> ⚠️ **PREPARE projects:** These use a Staging Folder (sandbox) exclusively. All content must be created within the Staging Folder — do not create filerooms or folders outside it. Use `listFolderContents` to locate the sandbox (type: SANDBOX, name: "Staging Folder") before creating any content.
**Two options — choose based on whether the project already exists:**
**Option A — New project (project does not yet exist):**
Use `listSubscriptions` to find the available subscription, then call `setupProject` with the confirmed folder tree as a `contentTree` JSON array. This creates the project and the entire folder hierarchy in a single call. Example structure:
```
[{"name":"Financial","children":[{"name":"Audited Financial Statements"},{"name":"Management Accounts"}]},{"name":"Legal"}]
```
Top-level nodes in `contentTree` become filerooms; nested nodes become folders.
**Option B — Project already exists:**
1. `listFolderContents` (no `metadataId`, `depth: 1`) — check if a fileroom already exists. Returns immediate top-level items only. If a fileroom exists, ask the user whether to add the index inside it or create a new one.
2. `createContent` — create folders inside the existing fileroom. Pass the full tree as `contentTree` with the fileroom's `metadataId` as `parentId`.
**Workflow:**
1. `listFolderContents` (no `metadataId`, `depth: 1`) — check if a fileroom already exists.
2. `createContent` — create the top-level fileroom if needed.
3. `createContent` — pass the full folder tree as `contentTree` so the entire hierarchy is created in one call.
**Error handling:** if a folder fails, note it and continue. Report failures at the end with the folder path so the user can investigate.
**On completion:**
> "Done ✓ — **[N] folders** created in **[Fileroom Name]**. Top-level sections: [list].
> 🔗 Open in Datasite: `https://app.global.datasite.com/en/platform/prepare/[projectId]/overview`
> Let me know if you’d like to adjust anything or invite team members."
**Important:** Always use the URL format above when linking to a Datasite project — `https://app.global.datasite.com/en/platform/prepare/{projectId}/overview`. Never construct a Datasite URL from memory or training knowledge; the format above is the only correct one.
---
## Reference materials
Read `references/sector-templates.md` for the full folder structures for each sector. Load only the section(s) relevant to the current deal — there's no need to read the whole file.
**Sectors covered:** Due Diligence (universal baseline), Technology, Healthcare, Healthcare Capital Raise, Manufacturing, Retail, Financial Services, Legal, Oil & Gas, Real Estate, Telecommunications, Transportation, Defence.
---
## Common Issues
**`getProjectOverview` fails or returns the wrong project**
Check that the Datasite MCP connector is connected (Settings → Extensions → Datasite should show "Connected"). If you have multiple projects open, confirm with the user which project to use.
**`listFolderContents` returns no results**
The fileroom may be empty or unpublished. Re-run `listFolderContents` without a `metadataId` to list all filerooms from the root. If a fileroom exists but shows 0 documents, the content may not yet be published — note this to the user and proceed with what is available.
**`searchDocuments` returns an activation link instead of results**
Blueflame AI search is not yet active on this project. Follow the Blueflame prompt in the skill instructions above. Do not attempt to answer using Claude's training knowledge.
**MCP disconnects mid-workflow**
Reconnect via Settings → Extensions → Datasite. Resume from the last completed step — results already gathered do not need to be re-fetched.
**`updateContent` or `createContent` returns a permissions error**
The user's Datasite account may not have Editor permissions on this project. Ask them to check their role in Datasite project settings.
Referenced files: 1
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- Datasite
Package observed Sep 30, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 18:00 UTC
- Collection status
- Collected
plugin_asdk_app_69eba17551ac81918231c83822b703b6
Download plugin data (JSON)