DEV Community

Cover image for FinePrint: an agent that checks your hackathon entry against the rules it reads
Himanshu Kumar
Himanshu Kumar Subscriber

Posted on Edited on AI-assisted

FinePrint: an agent that checks your hackathon entry against the rules it reads

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content.

What I Built

FinePrint checks a saved hackathon entry as the rules and your project facts change. When this challenge changed its entry-limit wording, the same two-entry plan went from Rules unclear to Blocked. The review keeps the old and new dated sources and identifies the finding that needs rechecking.

Sanity Context caught the rule change before I noticed it. After I added the official DEV pages to the Knowledge Base, Context flagged my September 20 "no limit" record against the revised one-submission-per-path wording. I reviewed the conflict, resolved it in favor of the live contest page and published the September 24 pack. Content Lake keeps both dated versions and their linked requirements, so FinePrint can trace the source change to a saved finding and recheck the same project facts.

Ask, “I am planning two entries in Path One.” The agent reads the Knowledge Base, quotes your fact and runs the typed check, showing the source reads and condition.

Try FinePrint · Open the guided demo · Watch the 65-second demo · Source code

The challenge page, FAQ and contest rules can disagree. I wanted an answer I could check, including “I still need a fact” and “these sources disagree.”

Each requirement links an event, a typed condition and a dated source. The model proposes facts and explains sources; the checker evaluates conditions. On September 20 the FAQ limited entries to one per path while the contest rules allowed unlimited entries. FinePrint returned Rules unclear and a question for the organizer. The September 24 pack, 2026-09-24.1, keeps both dated quotes and makes the same two-entry declaration Blocked.

GIBC V2’s Open Invention track is the second curated event, with 18 requirements alongside Sanity’s 19. Compare events carries team size and development date across both, leaving event-specific answers unknown: what existed before September 18 cannot establish what existed before July 11. Findings can be Supported, Blocked, Missing fact, Rules unclear or Not applicable. Reviews autosave in the browser and export as Markdown; a facts-only backup moves answers between browsers without importing a verdict.

You can also import a public rules page. The model proposes requirements, but each must quote the page word for word or be dropped and listed. FinePrint maps quotes onto a fixed fact vocabulary and builds the conditions itself. Unmapped rules and numbers absent from the quote become Check yourself, which can never show as Supported. Imports carry the real host and an “unreviewed” label, and stay in your browser.

For your own code, review a public GitHub repository against three curated judging rubrics: Sanity Path One, Path Two or GIBC Open Invention. FinePrint fixes the review to the current default-branch commit, reads up to ten files and 60,000 characters of selected text, and checks citations against inspected lines. Selected public code goes to Modal, never the public Sanity dataset. It runs no code and gives evidence, gaps and next steps; runtime behavior and judging scores remain outside this review.

Saved reports retain their requirements and sources. New curated packs identify added, removed or changed rules and affected findings. Capture-date refreshes alone do not change conditions. FinePrint checks curated Sanity records; it does not monitor organizer websites.

Demo

Start with the guided demo. It follows Two entries, one path without spending the live-agent allowance and labels its recorded evidence.

Inspect the entry-limit requirement through three states:

Selected rules and declared fact Entry-limit finding
September 20 pack; two entries in Path One Rules unclear: the FAQ and contest rules disagree
September 24 pack; the same two entries Blocked: one entry per path is now stated in both sources
September 24 pack; change the declaration to one entry Supported for this requirement; any other requirements with missing facts stay visible

The dated packs are stored in Sanity. The guide replays this sequence locally; its optional live question makes a fresh Context and model request. Changing the declaration does not edit either rule pack or certify the whole submission.

For a live run, open the review desk, create a Sanity review and use Ask FinePrint at the top:

  1. Choose Two entries, one path and press Ask FinePrint.
  2. Expand the trace to see initial_context, the Knowledge Base reads and check_requirements.
  3. Inspect the quoted fact, the cited September 24 entry-limit rule and the typed Blocked finding. Add these facts to my review preserves your other answers and notes.

The manual review desk remains available if the shared live-agent allowance is busy. Sample and recorded runs are labeled; they do not start a live model request.

A 65-second run on the deployed app, recorded September 25 with nothing sped up: the dated rule change, the agent answering a two-entries question through Sanity Context and an imported Devpost rules page.

FinePrint compares an illustrative five-person project against the Sanity and GIBC rule packs. Team size and the August start date are blocked for Sanity and supported for GIBC.

Browser capture from September 22. The example also declares that every member is a student. Both reviews retain unanswered questions.

Create a Sanity review, enter a team size of five and a development date of August 23, 2026, then open Compare events. Both checks are blocked for Sanity and supported for GIBC. GIBC still needs answers to its other requirements. Create a separate GIBC review to continue.

Open Try a rule-change rehearsal to lower a hypothetical team limit from six to three. One of the eighteen checks is affected. The rehearsal is labeled and local; it never changes an official source or saved review.

To check another event, start a review and pick Another hackathon (paste its rules link). Try https://zero-origin.devpost.com/rules, then press Read the rules. You get a review of that event's requirements, with the ones FinePrint could not map listed as Check yourself. Ask FinePrint works on it too.

The homepage labels its recorded source explanation and four comparison runs. Expand their source reads to inspect the Knowledge Base entries and structured records. The GIBC date example preserves a model/checker disagreement; the prior-work example lets you change whether an application or only its components existed before the event.

Live requests share a five-run allowance per ten minutes in a Vercel regional bucket, with additional checks that can limit repeated requests from the same network. The question and any form facts explicitly included with it go to the model provider. Personal reviews stay out of the public Sanity dataset.

Code

FinePrint uses Next.js, TypeScript, Zod, Sanity Content Lake and Context MCP. DeepSeek V4.1 Flash runs on an existing Modal endpoint. Vercel hosts the app.

Start with question-agent.ts for the tool loop, question-facts.ts for quote validation and engine.ts for the checks.

Tests cover review migration, event isolation, unknown facts, source changes, rule additions/removals, deadlines and bounded requests. The engine matches 42 authored scenarios; an earlier model baseline given the then-current nine Knowledge Base entries matches 39. I authored the labels and pack, so this measures agreement, not independent accuracy or retrieval quality.

How I Used Sanity

The public production dataset in project cxbqxkq6 contains two competitions whose current packs hold 37 requirements, plus the dated source versions each pack quotes. Superseded packs stay in the dataset. Each requirement references its event’s official sources. Event IDs, pack versions and explicit source references keep identical concepts such as team size separate. A GROQ query loads the pack, and Zod validates it before the checker uses it.

The content model:

Competition → versioned Requirement → dated SourceVersion
Knowledge Base → Context reads → quoted project facts
Validated facts + Requirement conditions → typed findings
Enter fullscreen mode Exit fullscreen mode

A saved finding records its requirement and source versions. Changing a team-size requirement identifies the team finding for review; refreshing a capture date leaves the condition unchanged. Content Lake supplies the records and relationships used for retrieval, comparisons and report updates.

Knowledge Base kbyrY7h8fTnL reads the public dataset through the dedicated fineprint MCP endpoint. Its purpose now covers Sanity and GIBC Open Invention, with separate event paths, exact rule titles and source versions. The app uses a server-side Context Viewer token. The browser never receives that credential.

For each question:

  1. Call initial_context to get the Knowledge Base outline.
  2. Select relevant paths and call knowledge_base_read. The server validates each path against the outline before reading it.
  3. Call check_requirements with quoted project facts and an independent interpretation of the relevant rules.
  4. Validate those proposals and run the typed conditions. Return the model’s interpretation alongside the checker’s result, including disagreements.
  5. Write the answer and cite retrieved entry paths. The question flow removes and discloses citation paths that were never read, and rejects an answer if no valid citations remain.

A live run on the deployed app: the agent reads the Sanity Context outline, three Knowledge Base entries and the structured requirements, then answers Blocked for two Path One entries with both its reading and the rule check shown

Live run on the deployed app, September 28: 5.7 seconds, eight recorded steps.

The question flow allows four model rounds, six entry reads and 45,000 source characters. The trace shows the actual calls and elapsed times. A provider error produces an error state with the completed steps.

The Knowledge Base reads four sources: the current curated production packs and three website sources, the challenge page, Sanity contest rules and official hackathon rules. Adding those pages exposed the stale entry-limit record described above. Context also flagged my “after the opening moment” wording against “during, and not prior to, the Entry Period”; the checker accepts the opening moment. I resolved both conflicts against the official sources and added a standing instruction dating the old entry-limit wording as a rule change.

For an imported event, the agent still reads the Knowledge Base for how FinePrint weighs sources, and uses a separate imported_rules_read tool over the quoted requirements. The trace names both, and citations are limited to what was actually read in that run.

What the live runs showed

The September 29 production check asked about a five-person team entering Path One. The agent read the Knowledge Base, accepted the quoted team-size fact and returned the correct Blocked finding. Its nine-step trace completed in about 6.3 seconds. A separate cited explanation completed in about 6.0 seconds, and a third request from the same network in that window returned the rate-limit message without starting a model call. The release passed 345 automated tests, the production build and 13 checks with providers disabled. Those checks cover different boundaries; a passing local test alone does not establish a working hosted agent.

The guided-demo browser check showed the three entry-limit states and completed its optional live question in 8.3 seconds: nine steps, four retrieved entries, the quoted fact and Blocked finding. Desktop and 390-pixel checks passed.

Earlier integration runs exposed disagreements. Four real Sanity Context/Modal calls checked team size and development date for both events. All typed checks matched my authored labels; the model agreed on three. Calls took 9.2–14.9 seconds and cited retrieved entries. On the GIBC date question the model requested a timezone despite August 23 being well inside the build window; the checker supported the date. The questions, answers and traces retain that failure. Four questions cannot establish general accuracy.

A September 22 local run against real services also kept a missing fact visible: “started in August” did not establish that the whole application already existed. The model called prior work blocked; the checker required that fact. A second run returned Rules unclear for two Path One entries against the September 20 pack. That historical result changes to Blocked under the September 24 pack.

I tested the importer on three real pages on September 24, before the stricter September 29 guards. These are historical counts, not a remeasurement of the current importer:

Page Rules found Mapped to a check Check yourself Quotes dropped
Zero Origin (Devpost rules) 18 4 14 0
This challenge's DEV page 16 8 8 0
A lablab.ai event plus its guidelines page 9 2 7 0

The first importer mapped “AI tools are permitted” to “AI use disclosed” and placed a start date nine hours early. I required timezone wording and words supporting the chosen fact mapping. A production Zero Origin import took 4.3 seconds; a local question against real services (“We are a team of five and the youngest of us is 16”) took 18.5 seconds over three rounds. Both agent and checker blocked team size.

The September 29 guards require a numeric bound and its direction, and a full quoted calendar date, time and timezone for date checks. Optional wording cannot become mandatory; a denial elsewhere in a sentence cannot negate an affirmative project fact. Ambiguous claims stay unresolved. Imported questions prefer server-cached rules; answers using only browser-supplied quotes disclose that they were not rechecked against the website.

JavaScript-rendered pages may expose little text: one lablab event returned about 1,000 characters and needed its guidelines page. Model mappings vary and remain unreviewed. These runs test particular questions and pages, not general eligibility accuracy.

Working with coding agents

I began with a pre-submission agent, chose “Revise the concept together before building” and narrowed it with “uncertain eligibility.” Codex built the first version using Modal, a public rules dataset and a read-only Context token. I requested a Three.js landing page, then chose Claude’s product-focused replacement and public question flow. Codex continued when Claude reached its session limit.

I pushed for multi-event comparison, rule-change impact and “all of the hacks if possible.” Claude built the importer; the live runs exposed its loose mappings. These additions made the boundaries concrete: team size can travel between events, event-specific declarations usually cannot. An English question cannot prove an English submission; a planned integration cannot prove a working one. Users inspect quoted facts before applying them.

Sanity Project Details

  • Project ID: cxbqxkq6, dataset production (public, read-only rule records)
  • Public dataset query: competitions with their requirement counts
  • Document types: competition, requirement, sourceVersion, reviewRubric
  • Knowledge Base: kbyrY7h8fTnL, served through the named Context MCP endpoint fineprint

No login is needed to try FinePrint. Rule checks run without the model. Live answers share a five-run allowance per ten minutes, with additional checks for repeated requests from the same network. A supported finding describes the selected rules and declared facts; the organizer makes the eligibility decision.

Agent Session

Building FinePrint with Codex: a rules agent over Sanity Context
Agent

I’ll read the attached request and follow its scope from there.

2 messages not shown
Agent

I’m using the hackathon-brainstorm skill to sharpen the FinePrint brief and verify the Sanity challenge rules. The workspace is empty, so I’m first checking the sponsor integration and the conflicting submission rules described in the attachment.

3 messages not shown
Agent

The submission-limit conflict is real: the FAQ allows one submission per path, while the contest rules allow unlimited entries. FinePrint fits Path One’s requirement for an agent using Sanity Context and a Knowledge Base.

I recommend the three-competition review desk in your brief, with the Sanity challenge as the first complete demo. The main tradeoff is curation effort: every check needs reviewed scope and evidence before it can produce a reliable result.

Your hackathon-brainstorm skill says, “Before scaffolding or implementing the product, present a recommended direction with its main tradeoff and ask the user to choose or revise it.” That is why I’m asking for a direction choice before building; I’ll continue the source checks meanwhile.

2 messages not shown
Agent

Let’s sharpen the concept before building. I see two weak points in the brief: maintaining three competitions could consume time that should go into the checking engine, and detecting conflicting sentences alone doesn’t prove that the agent understands which rules apply.

The stronger promise may be: “Show me what could invalidate my submission, and what changes when I fix it.” I’ll compare that with the track-and-prize focus while we choose the main problem.

3 messages not shown
Agent

I found close overlap with existing projects: ProofFlow describes rules-to-evidence audits, and SubmitGuard describes submission checks and corrective actions. Those are their published descriptions; I haven’t tested them.

That changes my recommendation: FinePrint needs a more specific advantage than producing a submission report. The strongest candidate is showing which rules apply to your project, why they apply, and how changing one fact changes the answer.

2 messages not shown
Agent

I’d keep FinePrint, but make “which rules apply to my situation?” its central problem.

The broad submission-auditor idea already has close neighbors. ProofFlow describes matching rules to project evidence; SubmitGuard describes checking submissions and fixing gaps. Those are published claims, not capabilities I tested.

There are three useful directions:

Direction Main benefit Main tradeoff
Submission repair desk Find missing evidence and prioritize fixes Considerable overlap with existing projects
Eligibility what-if desk Explain how project facts change applicable requirements Requires careful rule modeling
Organizer rules audit Find contradictions before participants encounter them Sanity already provides conflict detection

My recommendation is the what-if desk, with submission readiness as its output.

The demo could use actual Sanity requirements:

  1. Select Path One. Show the requirements applicable to that path.
  2. Change “requires login” to “yes.” Testing credentials or instructions become relevant.
  3. Add a second submission to the same path. Surface the disagreement between the FAQ’s limit and the contest rules’ unlimited-entry statement. Challenge requirements, contest rules.

The useful distinction is whether an ambiguity affects this particular submission. A conflict can exist without blocking every participant.

I’d also revise the scope:

  • Start with one deeply modeled competition, then add a second to test whether the approach generalizes.
  • Ask only for facts relevant to the selected path.
  • Distinguish what the builder declares from what the app actually verifies.
  • Preserve unresolved claims explicitly. Sanity’s conflict-resolution workflow asks someone to choose a claim; FinePrint should not make that decision without supporting authority. Sanity documentation.

Which has hurt you most in practice: overlooked evidence, uncertain eligibility, or contradictory rules? That should determine which direction we take.

Then I’d focus FinePrint on eligibility decisions made before someone commits to a build: whether their existing work, team, and intended track fit the rules. I’ll check the prior-work clauses closely, since they’re a useful test of whether the app can explain a decision without inventing certainty.

1 message not shown
Agent

Then FinePrint should help you decide whether to commit to an event with a particular project. Eligibility becomes the main workflow; submission artifacts come later.

A builder starts with a concrete question:

“I already have an app. Can I extend it for this competition, and which tracks could it qualify for?”

FinePrint asks only the missing questions that could change the answer, then produces separate findings for entering the event, qualifying for a prize, and receiving multiple prizes.

Each finding should say:

  • Supported: the supplied facts satisfy the applicable rule.
  • Blocked: a specific fact violates a requirement.
  • Missing fact: the answer depends on something the builder hasn’t supplied.
  • Organizer clarification needed: the sources don’t support a definitive answer.

There’s a better demo here than the submission-limit contradiction. Sanity’s FAQ encourages substantial improvements to previous work, while DEV’s general rules require development of the entry to begin during the contest period. Those statements aren’t necessarily contradictory: reusing a component in a new entry differs from entering an existing application. FinePrint needs to identify that distinction and ask what actually existed beforehand. Challenge FAQ, general rules.

I’d make that the first scenario: compare “new project using existing components” with “existing project receiving new features.” Show which findings change, the evidence behind them, and any interpretation still needing an organizer.

Sanity would hold the linked rules, their scope, source versions, and reviewed interpretations. The agent retrieves that context; explicit checks evaluate the project facts. The hard part—and the useful part—is deciding whether a rule applies.

The tradeoff is that some cases must remain unresolved. FinePrint should make the remaining uncertainty precise enough to act on.

What’s one real eligibility question you’ve been stuck on? A rough example is enough to shape the first workflow.

I’ll choose the eligibility scenario and check whether FinePrint is a strong fit for the Sanity challenge. If it holds up, I’ll proceed with the build; if the fit is weak, I’ll explain why and recommend a replacement first.

1 message not shown
Agent

FinePrint fits Path One: the challenge explicitly asks for an agent whose answers depend on structured content served through Sanity Context and a Knowledge Base. I’ll keep it and use existing-project eligibility as the main demo. Challenge brief

I’m using Impeccable for the interface: a focused review desk with project facts, eligibility findings, and the source behind each finding.

4 messages not shown
Agent

The key distinction is now clear: reusing a component and entering an application that already existed are different cases. I’m building the checks around that distinction, with explicit “missing fact” and “rules unclear” results.

I’m also checking Sanity access now, because a working Knowledge Base connection is required for the Path One submission.

11 messages not shown
Agent

Thanks. I’ll connect the curated rule set to Sanity next. The checking logic is taking shape: an existing application produces a development-window blocker, while a new entry using credited components is assessed separately. The examples are labeled as fictional declarations.

8 messages not shown
Agent

I’m building the review screen now. Changing a project fact will show which findings changed, and saved snapshots will let you compare two versions of a plan. Each finding will keep its scope visible, so a prize restriction isn’t mistaken for a ban on entering.

3 messages not shown
Agent

The OpenAI API key skill requires: “Before creating a key or writing any secret, obtain explicit confirmation.” That’s why I’m asking about the model connection. I’ll finish the independent rule checks and interface verification meanwhile.

5 messages not shown
Agent

The first 50 tests pass, covering unknown facts, opening-time boundaries, team limits, prior work, and the difference between entry and prize restrictions.

Sanity sign-in succeeded, but its Manage page still isn’t displaying projects or organizations. I’m checking the CLI path as well while I bring up the app for browser testing.

3 messages not shown
Agent

I’ll use Modal for the model and keep Sanity as the source of rules and Knowledge Base context. I’ll check the existing Modal setup and wire the agent to it.

4 messages not shown
Agent

I’m applying the Modal skill to use your existing workspace. The rule engine already passes 50 checks; I’m now connecting the source explanation flow and tightening the app’s saved reports and date handling.

7 messages not shown
Agent

Sanity CLI sign-in succeeded, and Modal authentication is working. I’m checking the available inference endpoint now. FinePrint will keep rule decisions deterministic; Modal will retrieve context and explain a finding without changing its status.

6 messages not shown
Agent

The Sanity rule records and seed script are ready for the new project. I’m continuing the Modal connection and app verification while that approval is pending.

Automatic approval review rejected creating the public Sanity project because the existing authorization covered connecting Sanity, but did not explicitly cover creating a persistent public dataset.

8 messages not shown
Agent

FinePrint’s Sanity project is created. I’m loading the official rule records next, then connecting the Knowledge Base. Modal’s existing DeepSeek endpoint is available, so I’m verifying that connection before provisioning anything new.

16 messages not shown
Agent

The live checks now work: Sanity returns all 19 requirements and three sources, and Modal answered a test about conflicting entry rules. I’m finishing the Knowledge Base retrieval path and testing the review flow on desktop and mobile.

6 messages not shown
Agent

The production build and all 60 tests pass. Browser verification hit a local disk-space error; I’m checking this project’s generated caches so I can clear safe, reproducible files and continue.

21 messages not shown
Agent

The disk issue is cleared, and the Knowledge Base has been created. Desktop and mobile layouts are working, including the conflicting-rules view. I’m checking saved-report restoration and linking the Knowledge Base to the agent.

29 messages not shown
Agent

The Knowledge Base is built. It flagged two equivalent deadline timestamps as a conflict; I’m reviewing that false alarm while keeping the real entry-limit disagreement visible. The independent UI review found two small readability fixes, which I’m applying.

Browser security rules require confirmation at the point of creating a new credential. The prepared token grants only Context Viewer access and expires after the challenge.

1 message not shown
Agent

The public rules dataset and Knowledge Base are created. I’m finishing the source-agent connection, applying two readability fixes from the review, and checking the full eligibility flow.

1 message not shown
Agent

I’ll connect the approved read-only token and use the Three.js scroll-worlds skill for FinePrint’s landing page. The landing page will introduce the eligibility workflow and lead into the review desk.

9 messages not shown
Agent

The read-only token is now stored locally. For the landing page, I’m building one continuous scene: a rulebook opens into separate requirements, conflicting rules come into view, and the pages assemble into a review. Native scrolling will control the scene, with readable text and a static fallback.

11 messages not shown
Agent

The landing page’s four-part scene is implemented. The core app’s 66 tests pass, including checks that reject invented source paths and citations. I’m now connecting the Knowledge Base endpoint so I can test the complete Modal-to-Sanity flow.

6 messages not shown
Agent

I’m continuing with the landing page and live source-agent check, then I’ll verify the mobile view and leave the app ready to try.

5 messages not shown
Agent

The FinePrint MCP endpoint was created. I’m testing it directly from the app now; the browser’s approval check timed out while opening its settings, so I’m using the read-only API to continue verification.

12 messages not shown
Agent

The landing page is rendering in the browser. The live source check hit a connection error before authentication completed; I’m testing that connection while checking the scroll scenes and responsive layout.

9 messages not shown
Agent

I found and fixed a timing bug that froze the 3D scene during chapter navigation. The scene now follows the scroll position correctly, and the still-view control works. Sanity has also confirmed the endpoint exposes the two required read-only tools.

4 messages not shown
Agent

The live connection works end to end: Modal retrieved entries from FinePrint’s Sanity Knowledge Base and explained both test cases. It preserved “rules unclear” for two entries and “blocked” for an existing app. The public dataset contains the curated rules; saved dossiers stay in browser storage.

10 messages not shown
Agent

The build passes. I’m checking the landing page across phone, tablet, and desktop sizes, including reverse navigation and keyboard controls. I’m also running the requested model baseline against the authored test cases; the results will be labeled as fixture agreement, not real-world eligibility accuracy.

17 messages not shown
Agent

The browser check caught an origin-validation bug: Next.js represented the local server differently from the browser URL, so the source request was rejected before reaching Modal. I’m fixing that check and adding a regression test.

4 messages not shown
Agent

The source agent now works from the app interface, with three Knowledge Base citations shown for the entry-limit conflict. All 70 tests pass. The recorded baseline matched 39 of 42 authored labels; its three differences concern missing facts and whether a rule applies.

11 messages not shown
Agent

I’m finishing the landing page and review desk. The final review found four small fixes: clearer mobile controls, visible source citations, and two typography adjustments. I’ll apply them, verify the fallback view, and leave the local preview ready.

14 messages not shown
Agent

The production build passes, and the static fallback preserves the story and review links. I’m checking the corrected layouts at desktop, tablet, and phone sizes, then verifying the new citation list with the live source agent.

17 messages not shown
Agent

The live agent now exposes all three cited Knowledge Base paths, and the list expands with the keyboard on mobile. The connection control also has an accessible name. I’m saving the final review evidence and documenting the design and verification limits.

14 messages not shown
Agent

The reviewer scored all four fixes resolved. The final animated scroll and reload checks also completed with no new browser errors. I’m finishing the handoff files; FinePrint is running locally with Sanity and Modal connected.

9 messages not shown
Agent

The app, build, and browser checks are complete. I’m finishing the design reference and verification record, including the camera settings and performance checks that still need a physical device.

6 messages not shown
Agent

You can use the local preview now. The final reference files will accompany it, covering the reusable design tokens, source-agent behavior, and recorded evaluation results.

3 messages not shown
Agent

FinePrint is ready locally with the Three.js landing page and live Sanity + Modal explanations.

Open preview · Build notes

70 tests and the production build pass. Desktop, tablet, mobile, and fallback views were checked. The documented fallback reviewer scored all four final fixes resolved.

<oai-mem-citation>
<citation_entries>
MEMORY.md:676-682|note=[Modal connection and cost control context]
</citation_entries>
<rollout_ids>
</rollout_ids>
</oai-mem-citation>

I’ll check what remains to make FinePrint a shareable hackathon demo, then continue with the next release step.

4 messages not shown
Agent

The main release gap is the live agent’s reliance on the local Modal login. I’m preparing hosted access for it, with a demo access code and a shared request limit so a public link can’t trigger unrestricted model usage.

10 messages not shown
Agent

The hosted access flow is implemented. Missing or failed rate-limit configuration now blocks a live agent call, and access sessions expire after two hours. I’m testing those boundaries and the unlock flow before preparing the deployment.

4 messages not shown
Agent

All 110 tests and the TypeScript check pass. I’m checking the production build and the demo access form next, then packaging the hosting setup.

10 messages not shown
Agent

The production build passed. It flagged a deprecated region setting, which I moved into Vercel’s configuration. The browser check now uses a temporary local demo code so I can test access without creating any cloud credentials.

13 messages not shown
Agent

The access form works on desktop and phone: incorrect codes are rejected, valid sessions survive a reload, and locking clears access. The hosting package also excludes local credentials and provider scripts from uploads.

9 messages not shown
Agent

The compiled production server passed seven more checks, including rejecting live-agent requests when authentication or rate-limit configuration is missing. The deployment files also passed a scan for configured credential values. The remaining hosting step is connecting the cloud credentials.

4 messages not shown
Agent

The hosting plan and deployment package are ready.

Your earlier approval covered storing the Sanity token locally. Hosting adds Vercel as a credential storage location and requires a persistent Modal token, so I need approval for those specific actions.

5 messages not shown
Agent

I’ll deploy FinePrint to Vercel, connect the approved credentials, activate the request limit, and test a live source explanation on the public URL.

7 messages not shown
Agent

Vercel’s CLI needs a fresh sign-in, and the browser is already signed into your account. I’m completing that connection before creating the project and credentials.

14 messages not shown
Agent

The Vercel project is created and the Modal token is stored securely. Vercel’s link command replaced the local environment file during its automatic import. I’m restoring that configuration before continuing.

18 messages not shown
Agent

The local configuration is restored. I recreated the FinePrint credential with the same Context Viewer permissions and October 20 expiry. The production credentials are now being saved to Vercel; next I’ll activate the request limit and verify the hosted app.

13 messages not shown
Agent

The public deployment built successfully at fineprint-kappa.vercel.app. I’m activating the Firewall rule now, then checking the full live explanation flow.

20 messages not shown
Agent

The hosted access flow works, and the Firewall allowed five verification requests before rejecting the sixth. A real source explanation is running now through Vercel, Modal, and the Sanity Knowledge Base.

7 messages not shown
Agent

The live explanation succeeded with three retrieved Knowledge Base citations. FinePrint kept the result as “Rules unclear,” as expected for the conflicting entry limits. I’m finishing the mobile and session checks and recording the deployment details.

15 messages not shown
Agent

FinePrint is live.

Verified the Three.js landing, live Sanity/Modal explanations, access controls, and desktop/mobile layouts. All 110 tests and the production build passed.

Private demo access code

<oai-mem-citation>
<citation_entries>
MEMORY.md:676-677|note=[Modal cost controls and resource isolation]
</citation_entries>
<rollout_ids>
</rollout_ids>
</oai-mem-citation>

I’ll find Claude’s recent work on FinePrint, check the project’s current state, and continue from where it left off.

3 messages not shown
Agent

I found Claude’s local session for FinePrint. The deployed app’s files appear unchanged since the last release; I’m checking Claude’s unfinished plan and review findings before continuing.

2 messages not shown
Agent

Claude stopped at its session limit while upgrading FinePrint in ~/n/fineprint. You had approved a product-focused landing without 3D, a public “Ask FinePrint” flow with visible source reads, formatting, and redeployment. I’ll finish that work here, using Impeccable for the interface and No AI Slop for the DEV draft.

7 messages not shown
Agent

The recovery is clear: Claude added source tracing and fact extraction, but hadn’t connected them to a working question flow or changed the landing. I’m finishing those pieces so the agent reads sources, turns stated facts into a review, and keeps disagreements with the rules checker visible.

7 messages not shown
Agent

The new question flow is now connected to Sanity retrieval and the rules checker. I’m replacing the landing’s 3D scene with readable product examples, including the real submission-limit conflict, then I’ll test the public flow and deploy it.

11 messages not shown
Agent

The first verification pass is clean: 131 tests pass, and the production build succeeds. The new agent preserves partial dates, rejects unsupported fact proposals, and shows its reasoning separately from the typed check. I’m testing real provider calls and desktop/mobile behavior next; Vercel’s saved login needs refreshing before deployment.

18 messages not shown
Agent

The first real agent run succeeded in 12.8 seconds. It read four Sanity entries, extracted three supported facts, and refused to treat “started in August” as proof that the application already existed. The UI also exposed where the model’s interpretation differed from the typed check.

14 messages not shown
Agent

The second live check also passed: FinePrint read the rules and kept the conflicting entry limits marked “unclear.” I’m finishing the visual review and documentation, then publishing and redeploying the upgrade Claude left unfinished.

8 messages not shown
Agent

The fresh visual review returned “ship” for desktop and phone. The landing now demonstrates the product, shows both conflicting rules, and makes clear that fixing one fact can leave another blocker. I’ve also updated the DEV draft to remove stale 3D references and unverified claims.

9 messages not shown
Agent

All 131 tests, formatting, type checks, and the production build pass. I’ve applied the public-agent setting you approved; the existing request limit remains in place. Next are the repository upload and a live check on Vercel.

10 messages not shown
Agent

The publication scan found no credentials in the staged files, and Vercel’s five-request limit is active. I’m waiting for the updated design documentation to finish so the repository describes the interface that will ship.

5 messages not shown
Agent

The upgrade is live at fineprint-kappa.vercel.app. I’m checking the hosted question flow and source explanations without a demo code, including whether they make real Sanity and Modal calls.

12 messages not shown
Agent

Hosted checks passed: the public question flow and per-finding explanation both made real source reads without a code. The browser run returned in 6.4 seconds, showed eight actual steps, and preserved the conflicting entry limits. I’m publishing the verified source now.

9 messages not shown
Agent

The source is published at github.com/himanshu748/fineprint. The hosted phone flow also passed: applying the answer carried over only the stated facts and highlighted the changed checks. I’m recording the final release details.

8 messages not shown
Agent

I checked Claude’s session history. It stopped at its limit during the FinePrint upgrade; I recovered and finished that work.

  • Replaced 3D with the product-focused landing you approved.
  • Added public questions with real Sanity reads, quoted facts, and visible agent steps.
  • Published and deployed it, keeping the request limit.

Live FinePrint · GitHub

Verified: 131 tests, production build, mobile layout, and live hosted calls. The DEV draft is updated; video/session evidence and final submission remain.

<oai-mem-citation>
<citation_entries>
MEMORY.md:1725-1729|note=[CLI recovery of Claude project history]
</citation_entries>
<rollout_ids>
01a08010-6959-7fd1-8c86-4559ea1497aa
</rollout_ids>
</oai-mem-citation>

I’ll check the full experience as a new user, then fix the gaps that make FinePrint feel like a demo instead of a usable product.

2 messages not shown
Agent

The main gap is the first visit: it opens a fictional project with prefilled answers, while saving and returning to your own work take extra effort. I’m changing that to personal reviews with autosave, clear next steps, and backup/restore. Examples and technical details will stay available, but outside the main workflow.

I’ll also make the supported event clear so people know which rules FinePrint can actually check.

5 messages not shown
Agent

Reviews will save on the user’s device, so people can return without creating an account. The main rule checks will remain usable when the AI request limit is reached. I’m adding validation and recovery paths so unfinished input, failed requests, or imports don’t erase existing work.

8 messages not shown
Agent

The personal review workflow is in place. I’m now checking autosave, returning to a review, and backup recovery so people can safely use it for their own projects. I’m also making sure the rule checks remain useful when AI is unavailable.

8 messages not shown
Agent

The app now starts with the user’s project, guides them to missing information, and saves their progress automatically. I’m testing two important safeguards: imports must preserve existing work, and AI suggestions must keep the user’s other answers and notes.

15 messages not shown
Agent

The browser checks passed for creating, reopening, archiving, restoring, and importing reviews. An unfinished date survived a reload and was correctly rejected when checked. A live AI answer also finished and stayed saved after I navigated away.

I’m finishing the mobile checks and updating the instructions before publishing this version.

9 messages not shown
Agent

The build and all 152 tests pass. The public interface now centers on “My reviews,” with examples and technical diagnostics moved into secondary areas. It clearly states the current limits: one supported event, storage in this browser, and optional AI with a shared allowance.

I’m checking the compiled app and the deployed workflow next.

8 messages not shown
Agent

The compiled app passed 10 additional checks, including manual checks with cloud providers disabled. The two-tab test also passed: when another tab changes a review, autosave pauses visibly instead of overwriting it.

I’m publishing the update now, then I’ll verify a fresh review on the public site.

11 messages not shown
Agent

The update is live. On the public site, I created a fresh review, checked its rules without an account, and received a cited AI answer in 5.2 seconds. Applying that answer kept the review’s other facts intact.

I’m verifying the final save/reopen step and leaving the app ready to use.

5 messages not shown
Agent

Updated and live: FinePrint.

People can now:

  • Create their own review without an account.
  • Autosave, resume, archive, and restore their work.
  • Export reports and transfer reviews through backups.
  • Follow guided next steps, inspect sources, and optionally ask AI.

Verified desktop/mobile flows, saved reviews, downloads, and live AI. 152 tests passed.

Currently supports the DEV × Sanity Challenge. Reviews stay in the browser; cloud sync isn’t included.

<oai-mem-citation>
<citation_entries>
MEMORY.md:154-154|note=[Verified hosted behavior separately from configuration]
</citation_entries>
<rollout_ids>
</rollout_ids>
</oai-mem-citation>

I’ll check the broken rules link and reassess FinePrint against the actual judging criteria and current entries. Then I’ll recommend the improvements most likely to strengthen it—more supported events may not be the most valuable change.

1 message not shown
Agent

I’m using the hackathon-brainstorm guidance to compare improvements against the rubric. FinePrint has a solid working foundation, but I wouldn’t call it a likely winner yet: several current entries also emphasize citations and conflicting sources. It needs a clearer, demonstrable advantage beyond those features.

4 messages not shown
Agent

The rules page is reachable; FinePrint opened it in a separate tab while leaving you on the landing page. I’m changing the source links to open directly in the current tab.

For the competition, my recommendation is a focused upgrade: compare the same project across a few curated events, show which reviews are affected when a rule changes, and demonstrate the advantage over ordinary search with recorded tests. That would make Sanity’s structured content central to the product.

1 message not shown
Agent

Your hackathon-brainstorm workflow says to “Wait for the choice” before implementing a major expansion. I’m checking that scope because multiple events change the data model and source-maintenance work.

7 messages not shown
Agent

The rules link now opens the official page in the same tab, and the local browser check passed. I’m deploying that fix. The main competitive gap is still proof: the demo needs to show a useful decision that depends on Sanity’s structured relationships.

7 messages not shown
Agent

The production build passed. My recommended demo is: one project, different results across a few events, then a rule update that identifies exactly which saved reviews need checking again. That would demonstrate both everyday usefulness and why the content model matters.

5 messages not shown
Agent

I’ll build that direction: a small set of verified events, one project compared across them, and rule-version changes that flag affected reviews. The contest-rules link fix is already live and verified.

10 messages not shown
Agent

I found a useful second event: GIBC V2 is currently open and permits teams of six, while Sanity’s limit is four. Its build window also starts earlier. I’m curating GIBC’s Open Invention track so the comparison can show meaningful differences, while keeping event-specific answers separate.

8 messages not shown
Agent

The core changes now keep each event’s rules and reports separate. Comparisons carry over facts such as team size and development date; answers whose meaning depends on the event stay unanswered. Rule-change detection will identify specific changed requirements and linked sources.

13 messages not shown
Agent

GIBC’s 18 selected checks are now in the public Sanity dataset. I’m refreshing the Knowledge Base and testing that old reviews still open, comparisons preserve unknown answers, and a change to one requirement flags only the checks it affects.

17 messages not shown
Agent

The comparison is working in the browser: the five-person August project fails Sanity’s team-size and timing checks but passes those two GIBC checks. GIBC still shows missing facts, rather than declaring the project eligible. The tests and production build passed; I’m checking mobile use and live source retrieval next.

9 messages not shown
Agent

The phone layout fits without horizontal scrolling, and creating a GIBC review preserves the shared facts while leaving its event-specific questions unanswered. I’m also applying the No AI Slop editing guidance to update the submission draft so it describes the new workflow and its limits accurately.

29 messages not shown
Agent

The draft now leads with the paired event test and explains exactly what Sanity’s content relationships do. It also keeps the model’s timezone mistake in the results. That makes the claims easier to judge.

The new phone flow is working: “Add this fact” focuses the right GIBC field, and changing the student declaration updates just that finding. I’m finishing the release checks before publishing the upgrade.

15 messages not shown
Agent

I’ll use Impeccable to tighten FinePrint’s design system and study 21st.dev for useful layout and interaction ideas. I’ll also compare verified winning posts with Sanity’s judging criteria, then improve the interface and submission around what FinePrint can actually demonstrate.

3 messages not shown
Agent

Sanity’s brief gives us a clear priority: show why structured content is necessary to the agent. FinePrint’s strongest evidence is the chain from event → requirement → source version → affected finding. I’ll make that chain easier to inspect, alongside a clearer comparison layout and consistent status labels. Judging criteria

The saved landing-page brief isn’t mapped to a source file, so I’m reading it directly and preserving its established visual direction.

12 messages not shown
Agent

I’m making the evidence part of the design: visitors will be able to switch between the recorded team-size and start-date checks, inspect the source reads, and see the model disagreement. That gives judges something concrete to evaluate while helping users understand the results.

10 messages not shown
Agent

The updated comparison is working, and all 179 tests pass. The submission draft now leads with the paired test and explains Sanity’s role through the actual content relationships and source reads. I’m checking the mobile layout and production build before publishing the update.

7 messages not shown
Agent

Desktop and mobile checks passed, including the expandable source evidence. The production smoke test also passed all 12 checks. I’m deploying now; the article remains an unpublished draft, with the model’s failed interpretation included in the results.

9 messages not shown
Agent

The update is live. Hosted checks confirmed both Sanity rule packs, the paired event comparison, and the submission screenshot. A fresh GIBC question completed through Sanity Context and Modal in 6.8 seconds, and the per-finding source explanation also passed.

4 messages not shown
Agent

I’m applying the hackathon-brainstorm workflow because arbitrary-event imports and repository reviews expand the agreed scope. It says to “wait for the choice” before implementation; the direction question is pending while I check feasibility.

My recommendation is a review with file-and-line evidence, missing evidence, and concrete fixes. Repository code can support a rubric assessment, but it cannot prove deployment quality, user adoption, or how judges will score originality.

2 messages not shown
Agent

The current Vercel site is live, and both event packs are loading from Sanity. FinePrint’s public GitHub repository is also accessible.

For broader coverage, I recommend this flow: import the event’s rules and rubric → review the extracted criteria → inspect a fixed GitHub commit → produce cited findings and prioritized fixes. If an event page blocks access or omits its rubric, the app should ask for the missing text rather than invent criteria. The main tradeoff is that imported events need source review; they cannot immediately carry the same confidence as curated packs.

Not every hackathon yet. FinePrint currently supports Sanity and GIBC V2. I verified that the Vercel site is live and both rule packs load.

I recommend adding:

  • Event import: official rules and judging rubric, with missing or conflicting criteria flagged.
  • GitHub review: inspect a specific commit and cite supporting files and lines.
  • Actionable findings: rubric coverage, missing evidence, and prioritized fixes—without pretending to predict judges’ scores.

The scope question is pending above. Unless you supply another example, I’ll validate the upgrade against FinePrint’s repository and the Sanity Challenge.

<oai-mem-citation>
<citation_entries>
MEMORY.md:195-196|note=[Preserve source uncertainty and verify hosted behavior]
</citation_entries>
<rollout_ids>
</rollout_ids>
</oai-mem-citation>

I’ll check the proposed GitHub-and-rubric review against the Sanity Challenge’s requirements and judging criteria, including whether Sanity would have a meaningful role.

1 message not shown
Agent

Yes—the GitHub rubric-review direction fits Sanity’s Path One. Its criteria explicitly reward structured content, Sanity Context, Knowledge Bases, technical quality, and usability. Official criteria

The strongest implementation would:

  • Store versioned event → rubric criterion → required evidence relationships in Sanity.
  • Retrieve the relevant criteria through Sanity Context.
  • Inspect a specific GitHub commit and link findings to files and lines.
  • Report evidence found, gaps, and fixes, keeping unverified runtime claims separate.
  • Recheck affected findings when the rubric or repository changes.

My recommendation is to prioritize GitHub review for the existing events, then add broader event imports. That gives a clearer demonstration of Sanity’s value and fewer unreliable ingestion paths.

The current product already uses real Context reads. GitHub analysis is still proposed, not implemented. The existing version is live on Vercel.

Submission still requires a published DEV post and Sanity project ID/public dataset link by October 4, 11:59 PM PDT. Concept fit does not establish personal eligibility or guarantee a win. Submission rules

<oai-mem-citation>
<citation_entries>
MEMORY.md:195-195|note=[Preserve uncertainty and separate source retrieval from eligibility decisions]
</citation_entries>
<rollout_ids>
</rollout_ids>
</oai-mem-citation>

I’ll build GitHub rubric reviews for the two supported events first, using a specific commit and file-and-line evidence. Then I’ll test the review against FinePrint’s own repository and deploy it to the existing Vercel site.

2 messages not shown
Agent

The useful distinction is between eligibility and judging evidence. I’ll keep the existing rule checks separate and add a repository review for each event’s published rubric. The review will show what code supports, what only the README claims, and what still needs a live demo or user evidence.

4 messages not shown
Agent

The review backend is in place. It resolves an immutable GitHub commit, limits which files it reads, and validates proposed citations against the retrieved lines. Documentation-only evidence cannot receive the strongest evidence label. I’ve also published the three rubric records to Sanity; next I’m connecting the review UI and testing a real run.

10 messages not shown
Agent

The first live run reached the model, but returned quotations longer than the report allowed. I’ve tightened the prompt and bounded the excerpts while retaining line verification. Sanity has detected all three new rubric records; I’m rebuilding the Knowledge Base before the next run.

5 messages not shown
Agent

The form now shows which files and lines were inspected, preserves the reviewed commit, and offers a downloadable report. I’ve added checks for invalid repository URLs, invented citations, and README claims mislabeled as implementation. A live review is running through the browser now.

7 messages not shown
Agent

The browser review completed successfully. It found implementation evidence for structured content, code quality, and Knowledge Base use, and marked usability as partial because it had not inspected the UI components. The report cited the inspected lines and kept runtime behavior unverified.

All 191 tests and 13 production checks pass. I’m waiting for Sanity’s rubric rebuild to finish, then I’ll verify the updated source reads and publish.

6 messages not shown
Agent

The tested code is pushed, and Vercel is building the release. The report also survives a browser reload. This release supports public repositories against Sanity Path One, Sanity Path Two, and GIBC Open Invention; arbitrary event imports remain outside its coverage.

5 messages not shown
Agent

The GitHub review page is live on Vercel. I’m checking real hosted reviews for all three rubrics now. Sanity’s Knowledge Base rebuild is still running; the app can already read the published rubric records directly from Content Lake.

4 messages not shown
Agent

All three hosted rubric reviews passed, taking about 6–8 seconds each. They read the newly indexed judging entries through Sanity Context. One proposed citation failed line verification and was discarded. The existing eligibility agent also still returns “Rules unclear” for the conflicting submission limits.

1 message not shown
Agent

Live on Vercel: Review your GitHub repository.

It now:

  • Reviews a fixed commit against Sanity Path One, Path Two, or GIBC Open Invention.
  • Shows verified file-and-line citations, evidence gaps, and next steps.
  • Separates README claims from implementation evidence.
  • Saves the latest report locally and supports downloading it.

Verified all three rubrics on production, plus 191 tests and 13 production checks.

Coverage is a bounded static review—not every hackathon or proof that the code runs.

Top comments (4)

Collapse
 
botsailorofficial profile image
BotSailor •

This is such a practical use of agents. Hackathon rules can be surprisingly easy to misread, especially when the FAQ, contest rules, and challenge page change at different times.

What I really like about FinePrint is that it doesn’t just give a confident “yes” or “no.” Showing the exact source, keeping outdated rule versions, and allowing “missing fact” or “rules unclear” as real outcomes makes the system feel much more trustworthy.

The model/checker disagreement examples are probably my favorite part too. Instead of hiding uncertainty, you’ve made it something the user can actually inspect. That feels like a much healthier way to build AI systems where a wrong assumption can have real consequences.

Really nice project—and honestly, I could see this pattern being useful far beyond hackathons. Rules, policies, eligibility criteria, compliance… there’s a lot of messy real-world text that could benefit from this kind of approach.

Collapse
 
himanshu_748 profile image
Himanshu Kumar •

Thanks. That was the point: when the FAQ and the contest rules disagreed, I wanted the answer to say so instead of guessing. Keeping the old rule versions is how FinePrint caught the "one entry per path" change on September 24.

Collapse
 
himanshu_748 profile image
Himanshu Kumar •

I built it for myself.

Collapse
 
aarishmansur profile image
Aarish mansur •

damn cool 🫡