Hackathon Projects / explained simply쉽게 설명 ← Playbook   Principles   Governance   Devpost ↗
Shipped at Hackathons

Eighteen things I built — in plain language.내가 만든 열여덟 가지 — 쉬운 말로.

Real projects, taken from idea to a working demo. No jargon — just what each one does, why it matters, and where to try it.아이디어에서 작동하는 데모까지 실제로 만든 프로젝트들. 전문 용어 없이, 각각 무엇을 하고 왜 중요한지, 그리고 어디서 볼 수 있는지.

Agentic Ops Control Center dashboard — live demo screenshot
Live demo · splunkhec2.streamlit.app
01
Agentic Ops Control Center: DataHub Guardrails에이전트 운영 관제 센터: DataHub 가드레일
An AI agent that checks the data catalog before it touches your data.데이터를 건드리기 전에 먼저 ‘데이터 카탈로그’를 확인하는 AI 에이전트.
The problem

AI agents are being handed the keys to company data — but they don't know what a data engineer knows: which datasets are outdated, restricted, or unreliable. They act confidently and blindly.AI 에이전트에게 회사 데이터 접근 권한이 주어지고 있지만, 정작 데이터 담당자가 아는 것 — 어떤 데이터가 낡았는지, 제한적인지, 믿을 수 없는지 — 은 모릅니다. 그래서 자신 있게, 그러나 눈먼 채로 움직입니다.

What it does

Before the agent uses any dataset, it asks a “library catalog” (DataHub): who owns this, is it deprecated, is the quality good? If the data is risky, the agent refuses and names who to contact. And every time it flags a problem, it writes a note back into the catalog — so there's always a record of what the AI did and why.에이전트가 데이터를 쓰기 전에 ‘도서관 카탈로그’(DataHub)에 물어봅니다: 이 데이터 주인이 누구인지, 폐기된 건 아닌지, 품질은 괜찮은지. 위험하면 사용을 거부하고 담당자를 알려줍니다. 그리고 문제를 발견할 때마다 카탈로그에 기록을 남깁니다 — AI가 무엇을 왜 했는지 항상 추적 가능하게.

Why it matters왜 중요한가

In regulated settings you must be able to prove what the AI accessed and why. This makes the catalog itself the audit trail.규제 환경에서는 AI가 무엇에 접근했고 왜 그랬는지 증명할 수 있어야 합니다. 이 프로젝트는 카탈로그 자체를 감사 기록으로 만듭니다.

AI agentsDataHubGovernanceAudit trail
02 · Supply chainA control tower for supply-chain disruption
02
Freight Desk Ops Control화물 운영 관제탑
A control tower that turns supply-chain disruptions into a clear impact map.공급망 혼란을 명확한 ‘영향 지도’로 바꿔주는 관제탑.
The problem

A single disruption — a port delay, a supplier failure — ripples through a supply chain invisibly. By the time you notice, several customers are already affected.항구 지연이나 공급업체 문제 하나가 공급망 전체로 보이지 않게 번집니다. 알아챌 때쯤이면 이미 여러 고객이 영향을 받은 뒤입니다.

What it does

Like an air-traffic control tower for shipments: when something goes wrong, it doesn't just send an alert — it traces exactly which shipments, customers, and downstream steps are affected, and proposes fixes for a human to approve.화물을 위한 항공 관제탑처럼: 문제가 생기면 단순 알림에 그치지 않고, 어떤 화물·고객·후속 단계가 영향을 받는지 정확히 추적하고, 사람이 승인할 수 있는 해결책을 제안합니다.

Why it matters왜 중요한가

It turns a vague “something broke” into “here's exactly what's affected and what to do about it” — with a person making the final call.막연한 ‘뭔가 잘못됐다’를 ‘정확히 무엇이 영향받고 무엇을 해야 하는지’로 바꿉니다 — 최종 판단은 사람이 하도록.

Supply chainAIDependency analysisHuman-approved
03 · Local & private100% local AI that observes and heals itself
03
Enterprise AI Without Leaks or Overspend유출도 과지출도 없는 기업용 AI
A 100% local AI agent that watches itself through Splunk and heals itself.100% 로컬에서 돌며 Splunk로 스스로를 관찰하고 스스로 고치는 AI 에이전트.
The problem

Cloud AI tools send your source code and data to servers you don't control, and hide what they cost until the bill arrives.클라우드 AI 도구는 당신의 소스 코드와 데이터를 통제할 수 없는 서버로 보내고, 비용은 청구서가 올 때까지 감춥니다.

What it does

A fully local AI assistant — nothing leaves your machine, so code and data stay private. It streams its own cost, speed, and errors into Splunk, and when something goes wrong — a cost spike, an error surge — it automatically switches to a cheaper or safer setting to fix itself.완전히 로컬에서 도는 AI 비서 — 아무것도 기기를 벗어나지 않아 코드와 데이터가 비공개로 유지됩니다. 자신의 비용·속도·오류를 Splunk로 흘려보내고, 비용 급증이나 오류 폭증 같은 문제가 생기면 더 저렴하거나 안전한 설정으로 자동 전환해 스스로 고칩니다.

Why it matters왜 중요한가

Privacy, self-healing, and no surprise bills — the three things enterprises fear most about AI, solved at once.프라이버시, 자가 치유, 그리고 예상치 못한 청구서 없음 — 기업이 AI에서 가장 두려워하는 세 가지를 한 번에 해결합니다.

Local LLMSplunkAuto-remediationCost control
04 · HealthcareCatch fraud before the money is paid out
04
Healthcare Fraud & Abuse Detection의료 청구 사기·남용 탐지
Catch prescription and claims fraud before the money is paid out.돈이 지급되기 전에 처방·청구 사기를 잡아냅니다.
The problem

Fraud, waste, and abuse in healthcare and insurance claims cost billions every year, and there are far too many claims for humans to review by hand.의료·보험 청구에서의 사기·낭비·남용은 매년 수십억 달러의 손실을 냅니다. 그런데 청구 건수가 너무 많아 사람이 일일이 검토할 수 없습니다.

What it does

It combines fixed rules (known fraud patterns, like impossible billing combinations) with AI reasoning that explains why a claim looks suspicious — flagging prescription fraud and billing abuse for a reviewer before payout.고정된 규칙(불가능한 청구 조합 같은 알려진 사기 패턴)과, 왜 그 청구가 의심스러운지 설명해주는 AI 추론을 결합합니다 — 지급 전에 처방 사기와 청구 남용을 검토자에게 표시해 줍니다.

Why it matters왜 중요한가

Catching fraud before payout — not chasing the money afterward — is what actually saves insurers billions.지급 후에 돈을 쫓는 게 아니라 지급 전에 사기를 잡는 것 — 이것이 보험사에게 실제로 수십억을 아껴줍니다.

Rule engineAI reasoningHealthcare analyticsFraud detection
05 · FinanceMarkets and the economy, made easy to see
05
MacroPulse
Makes the stock market and the economy easy to see — and see the risk in.주식 시장과 경제를 한눈에 보이게 — 그리고 위험까지 보이게.
The problem

Economic and market data is overwhelming and scattered across dozens of sources — hard to turn into a clear picture of where things stand.경제·시장 데이터는 방대하고 수십 개의 출처에 흩어져 있어, 지금 상황이 어떤지 명확한 그림으로 만들기 어렵습니다.

What it does

An investment dashboard that visualizes markets and economic indicators together and flags risk — turning a wall of numbers into a picture you can actually read.시장과 경제 지표를 함께 시각화하고 위험을 표시해주는 투자 대시보드 — 숫자의 벽을 실제로 읽을 수 있는 그림으로 바꿔줍니다.

Why it matters왜 중요한가

You see the whole landscape — and the risk in it — at a glance, instead of piecing it together from scattered charts.흩어진 차트를 짜맞추는 대신, 전체 지형과 그 안의 위험을 한눈에 봅니다.

PythonData vizRisk analysisFinance
06 · Healthcare AITurn payer policy changes into a ranked action list
06
Payer & Clinical Intelligence Agents지불자·임상 인텔리전스 에이전트
Two governed agents that read payer policy changes and rank what to act on — with a human deciding.지불자 정책 변경을 읽고 무엇부터 처리할지 순위를 매기는 두 개의 통제된 에이전트 — 판단은 사람이.
The problem

Insurance (payer) policies change constantly, and a single change quietly affects thousands of claims. By the time a team notices, denials and disputes are already piling up — and no one can see exactly what changed or what it costs.보험사(지불자) 정책은 끊임없이 바뀌고, 변경 하나가 수천 건의 청구에 조용히 영향을 미칩니다. 팀이 알아챌 때쯤이면 이미 거절과 분쟁이 쌓이는데 — 정확히 무엇이 바뀌었고 얼마의 비용이 걸렸는지 아무도 못 봅니다.

What it does

The Payer Intelligence agent diffs each new policy against the old one to surface what changed, answers questions grounded in the actual text with citations (and says “not covered” rather than guessing), and ranks the dollar impact on real claims. The Clinical & Growth agent identifies patient cohorts in member-months and flags quality and service-line opportunities. Every consequential action waits in a human-approval queue; every read is logged with its evidence. By default it runs framework-free and fully offline — so every answer is reproducible and auditable — and a LangChain + LLM path can be switched on to phrase answers more fluently, without ever changing the cite-it-or-abstain rule.지불자 인텔리전스 에이전트는 새 정책을 이전 버전과 비교해 무엇이 바뀌었는지 짚고, 실제 조문에 근거해 출처와 함께 답하며(모르면 지어내지 않고 ‘해당 없음’이라 말합니다), 실제 청구에 미치는 금액 영향을 순위로 보여줍니다. 임상·성장 에이전트는 환자 코호트를 member-months로 식별하고 품질·서비스라인 기회를 표시합니다. 모든 중대한 조치는 사람 승인 큐에서 대기하고, 모든 조회는 근거와 함께 기록됩니다. 기본은 프레임워크 없이 완전히 오프라인으로 돌아 모든 답이 재현·감사 가능하고, 필요하면 LangChain + LLM 경로를 켜서 답을 더 자연스럽게 다듬을 수 있습니다 — 출처를 달거나 모르면 유보하는 규칙은 그대로 둔 채.

Why it matters왜 중요한가

In a regulated payer setting you must prove what the AI concluded and why. Citations, abstention, and human approval make this auditable, not autonomous — the same shape as a production payer-intelligence platform.규제받는 지불자 환경에서는 AI가 무엇을 왜 결론 내렸는지 증명할 수 있어야 합니다. 출처·유보·사람 승인이 이걸 자율이 아니라 감사 가능하게 만듭니다 — 실제 지불자 인텔리전스 플랫폼과 같은 형태입니다.

RAG + citationsPolicy diffLangChain + LLM (optional)Human-approvedHealthcare
07 · Data scienceA model you can argue with, not just read a score from
07
California Housing Atlas캘리포니아 주택가격 아틀라스
A textbook machine-learning dataset, turned into something you can actually interrogate.머신러닝 교과서 데이터셋을, 직접 캐물을 수 있는 도구로.
The problem

Model results usually arrive as one number — “error: $50,462” — which tells you nothing about where the model is wrong, or whether to trust it for the decision in front of you. Meanwhile the dataset's own defects stay buried in a notebook cell nobody reads.모델 결과는 보통 숫자 하나로 옵니다 — ‘오차 $50,462’. 이 숫자는 모델이 어디서 틀리는지, 지금 내리려는 판단에 믿고 써도 되는지 아무것도 알려주지 않습니다. 그러는 사이 데이터 자체의 결함은 아무도 읽지 않는 노트북 셀에 묻힙니다.

What it does

All 20,640 census districts sit on a map you can filter by income, price, and age. Any correlation can be clicked open into the scatter plot behind it. A 24-tree random forest runs inside the browser, so an estimate appears as you move the sliders — and it is always shown next to the actual values of comparable real districts, never on its own. Three models are benchmarked side by side, including a decision tree with $0 training error and $73k held-out error: overfitting you can see rather than be told about.인구조사 구역 20,640개 전부가 지도 위에 있고, 소득·가격·연식으로 걸러볼 수 있습니다. 상관계수는 아무거나 눌러 그 뒤의 산점도를 펼칠 수 있습니다. 24그루 랜덤포레스트가 브라우저 안에서 돌아 슬라이더를 움직이는 즉시 추정값이 나오고, 그 값은 항상 비슷한 조건의 실제 구역들 옆에 나란히 놓입니다 — 혼자 덩그러니 있지 않습니다. 모델 세 개를 나란히 비교하는데, 그중 결정트리는 학습 오차 $0에 실전 오차 $73k입니다: 과적합을 설명으로 듣는 게 아니라 눈으로 봅니다.

Why it matters왜 중요한가

The uncomfortable parts are the point. Prices in this 1990 census were truncated at $500,001, so any model trained on it under-predicts expensive housing — the page says so up front, flags estimates that drift toward that ceiling, and shows the damage in the error distribution. A model that admits where it fails is worth more than one that only reports its score.불편한 부분이 핵심입니다. 1990년 인구조사의 집값은 $500,001에서 잘려 있어서, 이 데이터로 학습한 모델은 비싼 주택을 반드시 과소예측합니다 — 페이지는 이 사실을 먼저 밝히고, 그 천장에 가까워지는 추정값에 경고를 붙이고, 오차 분포에서 그 피해를 보여줍니다. 어디서 틀리는지 인정하는 모델이 점수만 보고하는 모델보다 쓸모 있습니다.

scikit-learnRandom forestIn-browser inferenceData vizModel honesty
08 · Market analysisOne record high, hiding two markets moving opposite ways
08
Georgia Housing Divide조지아 주택시장의 분열
A statewide index at an all-time high, while 60 of its 157 counties are already falling.주 전체 지수는 사상 최고치인데, 157개 카운티 중 60개는 이미 하락 중.
The problem

Housing gets reported as one number per state or metro. But an average can sit at a record high while a large share of the places inside it are already falling — and if you only ever see the average, you cannot tell which of those two things is happening to you.주택시장은 보통 지역당 숫자 하나로 보도됩니다. 그런데 평균은 그 안의 상당수가 이미 하락 중이어도 사상 최고치일 수 있습니다 — 평균만 보면, 내가 그 둘 중 어느 쪽에 있는지 알 수 없습니다.

What it does

It maps all 157 indexed Georgia counties, shaded by gain, by how far each has fallen from its own peak, by price level, or by the month it topped out. Every county's path is drawn from December 2019 so the split is visible rather than asserted, and a peak-timing chart shows this is genuinely two distributions: 45 counties peaked in 2024 and drifted down, 86 set their high this year. A conditions panel compares June 2019 with June 2026 — same month, so seasonality is held constant.지수가 산출되는 조지아 157개 카운티 전체를 지도에 올리고, 상승률·자기 고점 대비 낙폭·가격대·고점 시점 중 원하는 기준으로 색을 입힙니다. 모든 카운티의 경로를 2019년 12월부터 그려 분열을 주장이 아니라 눈으로 보게 하고, 고점 시점 분포는 이것이 실제로 두 개의 분포임을 보여줍니다: 45개는 2024년에 정점을 찍고 흘러내렸고, 86개는 올해 최고치를 세웠습니다. 시장 여건은 6월 대 6월로 비교해 계절성을 고정했습니다.

Why it matters왜 중요한가

The counties falling fastest are the expensive, urban ones. Atlanta's five core counties gained the least of any group (+48.4%) and every one of them is past its peak, while rural counties nearly doubled and are still at their highs. Averages hide inversions like that — and decisions get made on averages.가장 빨리 떨어지는 곳이 가장 비싸고 도시적인 카운티들입니다. 애틀랜타 코어 5개 카운티는 모든 그룹 중 가장 적게 올랐고(+48.4%) 5개 전부가 고점을 지났습니다. 반면 시골 카운티들은 거의 두 배가 되고도 여전히 최고치입니다. 평균은 이런 역전을 가리는데 — 의사결정은 평균으로 이뤄집니다.

Zillow research dataChoroplethTime seriesPublic dataMarket analysis
09 · Healthcare AISeven agents for a hospital, and not one of them decides anything
09
Infection Prevention AI병원 감염관리 AI
An infection prevention platform for the hospitals that have half a person to run one.감염관리를 0.5명으로 굴리는 병원을 위한 감염관리 플랫폼.
The problem

A community hospital runs infection prevention on half a person to two people, and finds cases by reading charts by hand. The bottleneck is not clinical judgment — it is the hours spent locating the cases that deserve judgment. The same one or two people also owe the CDC their surveillance data, the pharmacy an antibiotic stewardship program, and an accreditor a survey-ready evidence trail.지역 병원은 감염관리를 0.5명에서 2명으로 운영하고, 차트를 손으로 읽어 사례를 찾습니다. 병목은 임상 판단이 아니라, 판단할 가치가 있는 사례를 찾는 데 들어가는 시간입니다. 그 한두 사람이 동시에 CDC 감시자료 제출, 항생제 스튜어드십 프로그램, 인증 실사에 낼 근거까지 책임집니다.

What it does

One skill defines the reasoning — evidence first, human approval, an audit record for every action — and seven agents inherit it. Surveillance raises candidate infections from lab, device, and admission data. Outbreak turns a cluster into a twelve-step investigation with an epidemic curve and hypotheses that must carry their own contradicting evidence. Stewardship builds a ranked daily worklist from fourteen flag types. Compliance separates a practice gap from a documentation gap, because the fixes are opposite. Reporting assembles the committee, board, and regulator packages from those four agents’ audit trail rather than recomputing anything — one frozen snapshot, one number for every audience, and an explicit restatement when a later reclassification changes a figure that already went out. Education triages the cause of a finding before writing a word of teaching material, and refuses to propose any when the real cause is an empty dispenser or a form with no field for the thing being asked — because “we re-educated the staff” is the documented non-fix that lets a system problem survive another year. The seventh agent deploys the rest: it profiles a specific hospital’s systems, feeds, and staffing, then computes which rules can honestly run there. It ships with a regression suite for the rule engines underneath: eleven synthetic boundary cases with expected verdicts, run in CI with no patient data and no engine required.스킬 하나가 사고방식을 정의하고 — 근거 우선, 사람 승인, 모든 행동에 감사 기록 — 일곱 개의 에이전트가 그것을 물려받습니다. 감시 에이전트는 검사·장치·입원 데이터에서 감염 후보를 올립니다. 발생 에이전트는 집단발생을 12단계 역학조사로 바꾸고, 유행곡선과 함께 각 가설이 스스로를 반박하는 근거까지 달고 다니게 합니다. 스튜어드십 에이전트는 14종의 플래그로 우선순위가 매겨진 일일 작업목록을 만듭니다. 규정준수 에이전트는 실무의 결함과 기록의 결함을 구분합니다 — 고치는 방법이 정반대이기 때문입니다. 보고 에이전트는 위원회·이사회·규제기관용 보고서를 앞선 네 에이전트의 감사 기록에서 조립합니다 — 숫자를 다시 계산하지 않고, 스냅샷 하나로 모든 대상에게 같은 숫자를 주며, 이미 배포된 수치가 나중에 재분류로 바뀌면 조용히 고치지 않고 정정 공지를 냅니다. 교육 에이전트는 교육자료를 한 줄 쓰기 전에 지적사항의 원인부터 가려내고, 진짜 원인이 비어 있는 소독제 디스펜서나 항목이 없는 서식이면 교육을 제안하지 않습니다 — ‘재교육 실시’야말로 시스템 문제를 1년 더 살려두는 가장 흔한 가짜 해결책이기 때문입니다. 일곱 번째 에이전트는 나머지를 배치합니다: 특정 병원의 시스템·데이터 피드·인력을 진단하고, 그곳에서 정직하게 돌아갈 수 있는 규칙이 무엇인지 계산합니다. 그 아래 규칙 엔진을 위한 회귀 테스트 세트가 함께 들어갑니다: 기대 판정을 붙인 합성 경계 사례 11건을, 환자 데이터도 엔진도 없이 CI에서 돌립니다.

Why it matters왜 중요한가

The design work is in what each agent refuses to do. Surveillance will not confirm a reportable event. Outbreak will not name a staff member or declare an outbreak. Stewardship will not phrase a flag as an instruction. Compliance has no individual-level output mode at all. And capability is derived from the data, never assumed: a rule whose required feed is missing is disabled and shown as disabled, so a hospital that can support 9 of 34 rules gets 9 rules and an honest list of the 25 it does not have. That claim gets tested rather than asserted: a bloodstream-infection engine written straight from the textbook definition passes all five typical cases in the suite and fails all six boundary ones — five over-counts and one missed newborn. Against a five-case test set it would have scored full marks and shipped.설계의 본체는 각 에이전트가 하지 않겠다고 정한 것들입니다. 감시 에이전트는 보고 대상 사례를 확정하지 않습니다. 발생 에이전트는 직원 이름을 적지 않고 유행을 선언하지도 않습니다. 스튜어드십 에이전트는 플래그를 지시문으로 쓰지 않습니다. 규정준수 에이전트는 개인 단위 출력 자체가 없습니다. 그리고 성능은 데이터에서 도출될 뿐 가정되지 않습니다: 필요한 데이터가 없는 규칙은 비활성화되고, 비활성화되었다고 표시됩니다. 그래서 34개 중 9개만 지원 가능한 병원은 9개와, 갖지 못한 25개의 정직한 목록을 함께 받습니다. 이 주장은 선언이 아니라 검증 대상입니다: 교과서 정의 그대로 짠 혈류감염 판정 엔진에 이 11건을 먹여보면, 전형적인 5건은 전부 맞히고 경계 6건은 전부 틀립니다 — 과다계수 5건과 놓친 신생아 1건. 5건짜리 테스트 세트였다면 만점을 받고 배포됐을 엔진입니다.

Agent architectureBoundary fixturesHAI surveillanceHuman-in-the-loopHealthcare
10 · Agent GovernanceFour agents, and the interesting part is the doors they cannot open
10
Gemini Ops Fleet제미나이 옵스 플릿
A fleet of back-office agents whose limits are enforced in code, not requested in a prompt.한계를 프롬프트로 부탁하지 않고 코드로 강제한 백오피스 에이전트 팀.
The problem

A fifty-person manufacturer with no IT department runs on four separate spreadsheets, and a support ticket that arrives at two in the morning waits until someone opens a laptop. Agents should help. But this is also the company that cannot survive one wrong email to a customer, or the margin memo reaching a sales rep before the quarter closes. Every agent demo answers “how much can it do on its own?” Nobody was answering “what is it structurally unable to do?” — which is the only question a company without a compliance team can actually act on.IT 부서가 없는 50인 제조회사는 영업·지원·회계·대표가 각자 엑셀을 씁니다. 새벽 2시에 들어온 문의는 아침에 누가 노트북을 열 때까지 그대로입니다. 에이전트가 도와야 할 회사죠. 그런데 동시에, 고객에게 잘못된 메일 한 통이 나가거나 분기 마감 전 마진 자료가 영업사원에게 넘어가면 버티지 못하는 회사이기도 합니다. 모든 에이전트 데모는 “혼자 얼마나 할 수 있나”에 답합니다. “구조적으로 무엇을 못 하나”에 답하는 곳은 없었습니다 — 컴플라이언스 팀이 없는 회사가 실제로 쓸 수 있는 답은 그것뿐인데도.

What it does

Four agents run off a shared event stream instead of a chat box: one classifies incoming tickets and assigns an owner, one turns resolved tickets into searchable documents, one drafts customer messages, and one reconciles the order book against the ledger, read-only. A business change writes its record and its event in the same transaction, so nobody is waiting at a prompt. Then three limits hold regardless of what anyone types. No tool accepts a role or an employee id, so the model has no vocabulary for claiming one — and a test asserts that across every tool signature, so the guarantee survives future changes rather than resting on discipline. A sales rep asking for the margin memo gets an empty result, not a refusal the model composed, because the document was excluded by a SQL filter before any row existed. And the drafting agent can only queue a draft: approving and sending are endpoints absent from every agent’s tool set, and send refuses outright if no human signed off.채팅창이 아니라 공용 이벤트 스트림에서 도는 에이전트 4개입니다. 하나는 들어온 티켓을 분류해 담당자를 정하고, 하나는 해결된 티켓을 검색 가능한 문서로 만들고, 하나는 고객 메시지 초안을 쓰고, 하나는 주문 장부와 회계 장부를 읽기 전용으로 대조합니다. 업무 변경은 기록과 이벤트를 같은 트랜잭션에 쓰기 때문에 누구도 프롬프트 앞에서 기다리지 않습니다. 그리고 누가 무엇을 입력하든 세 가지 한계가 유지됩니다. 어떤 도구도 직급이나 사번을 인자로 받지 않습니다 — 모델에게는 신분을 주장할 어휘 자체가 없고, 그 사실을 테스트가 모든 도구 서명에 걸쳐 검사하므로 사람의 성실함이 아니라 구조가 보장을 지킵니다. 영업사원이 마진 메모를 요청하면 빈 결과가 돌아옵니다 — 모델이 지어낸 거절 문구가 아니라, 그 문서가 한 행이 존재하기도 전에 SQL 필터에서 빠졌기 때문입니다. 초안 작성 에이전트는 초안을 큐에 넣는 것이 전부입니다: 승인과 발송은 어떤 에이전트의 도구 목록에도 없는 엔드포인트이고, 사람의 서명이 없으면 발송은 거부됩니다.

Why it matters

A guarantee the model is asked for is not a guarantee. The restrictions are written into every agent’s instruction and enforced in Python — and when the injection attempt was tested, the instruction turned out to be irrelevant: the guardrail stopped it before the model was called, and the tool permissions would have stopped it after. The instruction is a courtesy; the code is the contract. Four of the eleven cuts in the demo end in a failure that is supposed to happen, refusals are recorded in telemetry with the same weight as successes, and the whole system — all 49 tests, including every access-control and human-gate assertion — runs on a laptop with no cloud account, so a reviewer can verify the claims before deciding whether to trust the demo.모델에게 부탁한 보장은 보장이 아닙니다. 제약은 모든 에이전트의 지시문에 적혀 있고 동시에 파이썬 코드로 강제되는데, 인젝션 공격을 실제로 시험해 보니 지시문은 아무 역할도 하지 않았습니다 — 가드레일이 모델을 부르기 전에 막았고, 통과했더라도 도구 권한이 막았을 것입니다. 지시문은 예의이고, 코드가 계약입니다. 데모 영상 11컷 중 4컷은 일어나야 마땅한 실패로 끝나고, 거부는 성공과 같은 무게로 관측 기록에 남으며, 시스템 전체가 — 접근 통제와 사람 승인 검증을 포함한 49개 테스트 전부가 — 클라우드 계정 없이 노트북에서 돌아갑니다. 심사자가 데모를 믿을지 결정하기 전에 주장을 직접 검증할 수 있습니다.

Gemini + ADKEvent-drivenAccess controlHuman-approvedCloud Run
11 · Governed RetrievalSame question, two correct answers, and the date decides which
11
Policy Copilot정책 코파일럿
Governed question-answering over versioned policy, where which version governs is arithmetic rather than something you ask a model.버전이 있는 정책에 대한 통제된 질의응답 — 어느 버전이 지배하는지는 모델에게 묻는 것이 아니라 계산하는 것입니다.
The problem

Operations teams answer questions by finding the clause that governs. But policy has versions. The rule that applied to a March transaction may have been replaced in May, and both texts still exist, both are retrievable, and only one of them is the answer. Ask a retrieval system and it hands back today’s clause — or worse, both, with nothing marking which is which. The usual response is that the model will reconcile them. It won’t. It will guess, fluently, with a citation attached. And access is not uniform: some clauses are restricted, so the same correct answer becomes a disclosure when the wrong person receives it.운영팀은 지배하는 조항을 찾아 질문에 답합니다. 그런데 정책에는 버전이 있습니다. 3월 거래에 적용된 규칙이 5월에 교체됐다면 두 텍스트가 모두 존재하고 모두 검색되지만, 정답은 하나뿐입니다. 검색 시스템에 물으면 오늘의 조항을 돌려주거나, 더 나쁘게는 둘 다 돌려줍니다 — 어느 쪽이 어느 쪽인지 표시 없이. 흔한 대답은 “모델이 조정할 것”입니다. 조정하지 않습니다. 유창하게, 인용까지 붙여서 추측합니다. 게다가 접근 권한은 균일하지 않습니다. 같은 정답도 받는 사람이 다르면 정보 유출이 됩니다.

What it does

Retrieval takes the date the question is about, not the date it is asked. Ask about a March representment and you get the 30-day clause that governed then, with a note that a later version took over on 1 May; ask about June and you get the 45-day clause and what it replaced. Supersession is settled before the model sees anything, because whether a clause governs is arithmetic over effective dates. The predicate order carries the rest: entitlement and effective date run first, because they are security — a clause the caller cannot read is gone before ranking ever happens. Relevance runs after, on already-permitted material, so it can only remove an answer and never add one. Entitlements carry an epoch, so an index built under permissions that have since been revoked refuses to serve rather than quietly answering.검색은 질문이 다루는 날짜를 기준으로 합니다, 질문한 날짜가 아니라. 3월 대표권을 물으면 그때 지배하던 30일 조항과 “5월 1일부터 이후 버전이 적용된다”는 표시를 함께 받고, 6월을 물으면 45일 조항과 그것이 무엇을 대체했는지를 받습니다. 승계는 모델이 무언가를 보기 전에 해소됩니다 — 어느 조항이 지배하는지는 유효기간에 대한 산술이기 때문입니다. 나머지는 술어의 순서가 담당합니다. 권한과 유효기간이 먼저 실행됩니다, 그것이 보안이기 때문에 — 호출자가 읽을 수 없는 조항은 순위 계산이 시작되기도 전에 사라집니다. 관련성은 그다음, 이미 허용된 자료에만 적용되므로 답을 없앨 수는 있어도 늘릴 수는 없습니다. 권한에는 epoch이 붙어 있어서, 이미 취소된 권한으로 만들어진 인덱스는 조용히 답하는 대신 서빙을 거부합니다.

Why it matters

The interesting part was a failure. The obvious way to make a retriever decline is a similarity floor — and measured on this corpus, “What is the capital of France?” scored 0.408, higher than two of the three genuine policy questions. Bag-of-words similarity is dominated by the words two sentences happen to share — what, is, the, of — so it measures sentence shape rather than aboutness, and no threshold separates the two sets. That non-separability is now itself a test: if a better embedder ever makes them separable, the build fails and the mechanism gets revisited instead of carried forward out of habit. The gate that does work weights terms by how distinctive they are, so “merchant” decides nothing and “representment” decides everything. All 58 tests run with no API key and no network, because an access-control claim a reviewer cannot check without the author’s credentials is a claim taken on faith.흥미로운 부분은 실패였습니다. 검색기가 거부하게 만드는 뻔한 방법은 유사도 하한인데, 이 코퍼스에서 측정해 보니 “프랑스의 수도는?”이 0.408로 진짜 정책 질문 셋 중 둘보다 높았습니다. bag-of-words 유사도는 두 문장이 우연히 공유하는 단어 — what, is, the, of — 에 지배되므로 문장이 무엇에 관한 것인지가 아니라 문장의 모양을 잽니다. 두 집합을 가르는 임계값은 존재하지 않습니다. 이 분리 불가능성 자체가 이제 테스트입니다: 나중에 더 나은 임베더로 분리가 가능해지면 빌드가 깨지고, 메커니즘을 관성으로 끌고 가는 대신 재검토하게 됩니다. 실제로 작동하는 게이트는 용어를 변별력에 따라 가중합니다 — “merchant”는 아무것도 결정하지 않고 “representment”가 전부를 결정합니다. 58개 테스트 전부가 API 키 없이, 네트워크 없이 실행됩니다. 리뷰어가 저자의 자격증명 없이는 검증할 수 없는 접근통제 주장은, 믿음으로 받아들여야 하는 주장이기 때문입니다.

As-of retrievalAccess controlPolicy versioningAbstentionOffline tests
12 · Scoping & EstimationThe same 15,000 documents cost seven weeks or over a year, and the client picks which
12
Document Scoping Estimator문서 정리 견적 도구
A scoping tool for the question every data project opens with — and that should never be answered with a single number.모든 데이터 프로젝트가 시작하는 그 질문을 위한 스코핑 도구 — 그리고 단일 숫자로 답해서는 안 되는 질문.
The problem

A client asks how long and how much it will take to organize fifteen thousand documents. Answer with a number and you have made a guess that the client hears as a commitment. Worse, “organize” is three different projects sharing one word: building a searchable index, classifying and extracting entities, or reading across documents for findings that are not in any single one — and each one costs roughly triple the one beneath it. Meanwhile the question that actually drives the bill, how accurate does this need to be, is usually the one nobody has put to the client at all. So the team quotes a discovery draft and is later asked for a verified record.클라이언트가 묻습니다: 문서 15,000건 정리에 얼마나 걸리고 얼마나 드나요? 숫자로 답하면 추측을 한 것이고, 클라이언트는 그것을 약속으로 듣습니다. 게다가 “정리”는 한 단어를 쓰는 세 개의 다른 프로젝트입니다 — 검색 가능한 색인 구축, 분류와 개체 추출, 그리고 어느 한 문서에도 없는 결론을 위해 문서들을 가로질러 읽는 것. 각 단계는 아래 단계의 대략 세 배입니다. 그런데 정작 청구서를 결정하는 질문 — 얼마나 정확해야 하는가 — 은 아무도 클라이언트에게 물어본 적이 없습니다. 그래서 팀은 탐색용 초안을 견적하고, 나중에 검증된 기록을 요구받습니다.

What it does

Three inputs, because the question is three questions. Scope; source condition — native digital or scanned, and at what OCR accuracy; and the accuracy the client actually needs. It returns a band rather than a number, produced by a 1.5–2.0× overhead multiplier, and the band is wide on purpose: before anyone has measured anything, a narrow range is a guess wearing a suit. Then it splits the drivers by what kind of problem they are. Accuracy target and scope level dominate the total and no pilot will ever discover them — they are the client's call, and they belong in the first conversation rather than the third. Pages per document, OCR accuracy and review minutes are unknowable from a meeting and knowable in two weeks, which is exactly what a sample is for. It sizes that sample at three percent, shows the range collapsing afterwards, and writes out the paragraph you would say aloud.입력은 셋입니다, 질문이 셋이기 때문에. 스코프, 원본 상태 — 네이티브 디지털인가 스캔인가, 그리고 OCR 정확도는 얼마인가 — 그리고 클라이언트가 실제로 필요한 정확도. 결과는 숫자가 아니라 밴드이고, 1.5–2.0× 오버헤드 배수에서 나옵니다. 밴드가 넓은 것은 의도입니다: 아무도 아무것도 측정하지 않은 시점에 좁은 범위는 정장을 입은 추측입니다. 그다음 동인들을 문제의 종류별로 나눕니다. 정확도 목표와 스코프 레벨은 총액을 지배하지만 어떤 파일럿도 알아내지 못합니다 — 클라이언트가 결정할 사안이고, 세 번째 미팅이 아니라 첫 미팅에 나와야 합니다. 문서당 페이지 수, OCR 정확도, 검토 분은 회의로는 알 수 없고 2주면 알 수 있습니다 — 표본은 정확히 그것을 위해 존재합니다. 도구는 그 표본을 3%로 산정하고, 이후 범위가 좁혀지는 모습을 보여주고, 소리 내어 말할 문단을 그대로 써 줍니다.

Why it matters

Two things fall out of running it. Model inference is about one percent of the modelled cost at this volume — clients arrive wanting to discuss token prices, and the money is in build time and in how many documents a person has to open. And pushing the accuracy target from ninety-five percent to 99.9 on a scanned corpus takes the same fifteen thousand documents from roughly seven weeks to well over a year, because the staffing changed rather than the technology. The separation between deciding and measuring is the part worth keeping: conflate them and a team runs a pilot to answer a question the client could have settled on day one, or quotes a programme around an accuracy target nobody ever set. Every rate in the tool is a placeholder you replace before it reaches a client — the structure is the product, not the constants.실행해 보면 두 가지가 드러납니다. 이 물량에서 모델 추론은 모델링된 비용의 약 1%입니다 — 클라이언트는 토큰 가격을 논하고 싶어 하며 오지만, 돈은 구축 시간과 사람이 열어봐야 하는 문서 수에 있습니다. 그리고 스캔 코퍼스에서 정확도 목표를 95%에서 99.9%로 올리면 같은 15,000건이 약 7주에서 1년 넘게로 갑니다 — 기술이 아니라 인력 구성이 바뀌었기 때문입니다. 남길 가치가 있는 것은 결정과 측정의 분리입니다: 이 둘을 섞으면 팀은 클라이언트가 첫날 정할 수 있었던 질문에 파일럿을 돌리거나, 아무도 정한 적 없는 정확도 목표 위에 프로그램을 견적합니다. 도구의 모든 단가는 클라이언트에게 보이기 전에 교체할 자리표시자입니다 — 상품은 구조이지 상수가 아닙니다.

EstimationRisk framingPilot designClient-facingNo dependencies
13 · Science CommunicationClarity is what makes a reader certain — including certain of things the study never showed
13
Explain Research논문 쉽게 풀기 스킬
A skill for turning a dense paper into plain language without letting the simplification overstate it — ordered so the guards come before the polish.어려운 논문을 쉬운 말로 옮기되 단순화가 과장으로 넘어가지 않게 하는 스킬 — 다듬기보다 방어막이 먼저 오도록 배치했습니다.
The problem

Two failures sit either side of explaining research. Explain too little and the reader leaves with nothing. Explain too smoothly and they leave certain of something the study did not show — and the clarity is what made them certain. The second failure looks like success, which is why it is the common one: asked to simplify, anyone will drop the hedge before they drop the finding, because the hedge reads as clutter. The output is fluent, confident and slightly false. The version this replaces was four sound techniques — analogy, three-part chunking, translating numbers, sealing misreadings — but it was written as an essay about one paper, it never asked whether the paper deserved explaining, and it had no way to check its own output against the source.연구를 설명할 때 실패는 양쪽에 있습니다. 너무 적게 설명하면 독자는 아무것도 얻지 못합니다. 너무 매끄럽게 설명하면 독자는 연구가 보여주지 않은 것을 확신한 채 떠나고 — 그 확신을 만든 것이 바로 명료함입니다. 두 번째 실패는 성공처럼 보이기 때문에 흔합니다. 쉽게 만들라는 요청을 받으면 누구든 발견보다 단서를 먼저 버립니다. 단서가 군더더기로 읽히기 때문입니다. 결과물은 유창하고 자신 있고 약간 거짓입니다. 이 스킬이 대체한 이전 버전은 네 가지 건전한 기법(유추·3단 분해·숫자 번역·오해 봉인)이었지만, 한 논문에 대한 에세이로 쓰여 있었고, 그 논문이 설명할 가치가 있는지 묻지 않았으며, 자기 출력물을 원문과 대조할 방법이 없었습니다.

What it does

Six steps, and three of them are new. Name the reader first — “general audience” leaves every downstream choice undetermined and gives no stopping condition; a named reader decides the analogy and tells you when you are done. Assess the source before agreeing to explain it, because explaining a weak study clearly is worse than not explaining it: your clarity lends it credit it did not earn. Every analogy must state where it breaks, on the argument that if you cannot name the break you do not understand the analogy well enough to use it. Numbers are translated into what changed, carrying their uncertainty. Misreadings are sealed as what a reader might now wrongly believe rather than hedged. And finally, round-trip it: restate the finding from your own simplification alone and compare it to the paper's claim. If yours is stronger, it drifted — at a sentence you can now find.여섯 단계이고 그중 셋이 새로 들어갔습니다. 독자를 먼저 지정합니다 — “일반 독자”는 이후 모든 선택을 미결로 남기고 멈출 조건도 주지 않습니다. 독자가 정해지면 비유가 결정되고 언제 끝났는지도 알 수 있습니다. 설명하기로 하기 전에 원문을 평가합니다. 부실한 연구를 명료하게 설명하는 것은 설명하지 않는 것보다 나쁘기 때문입니다 — 명료함이 그 연구가 얻지 못한 신뢰를 빌려줍니다. 모든 비유는 깨지는 지점을 명시해야 합니다. 깨지는 지점을 말할 수 없다면 그 비유를 쓸 만큼 이해하지 못한 것이라는 논리입니다. 숫자는 불확실성을 달고 “무엇이 달라졌는가”로 번역합니다. 오해는 얼버무리는 대신 독자가 지금 잘못 믿게 될 것으로 못 박습니다. 마지막으로 왕복 검증: 자기가 쓴 쉬운 버전만 읽고 발견을 복원해 원문 주장과 비교합니다. 내 쪽이 더 강하면 표류한 것이고, 이제 어느 문장에서 표류했는지 찾을 수 있습니다.

Why it matters

It was tested on a Lancet Global Health study estimating that unilateral sanctions are associated with around 560,000 deaths a year, and the new steps earned their place immediately. The round-trip check caught a draft headline — “sanctions kill as many people as war” — that was stronger than the source on three counts at once: causal where the paper is careful, absolute where the estimate spans 370,000 to 760,000, and comparative in a way the authors did not claim. The source-assessment step surfaced a post-publication correction and a set of declared interests that the four-technique version had no step capable of finding. That test then produced a further fix: step five now says how to present declared interests — state them as fact, note that they were disclosed, put the study's strengths in the same breath, and never let an interest carry a criticism of the finding, because “the funder has a position on this” is not an argument against a coefficient.일방 제재가 연 약 56만 명의 사망과 연관된다고 추정한 Lancet Global Health 논문으로 시험했고, 새 단계들이 곧바로 값을 했습니다. 왕복 검증이 초안 제목 — “제재는 전쟁만큼 사람을 죽인다” — 을 잡아냈습니다. 이 문장은 원문보다 세 방향으로 동시에 강했습니다: 논문이 조심스러운 지점에서 단정적이고, 추정 범위가 37만~76만인데 절대적이며, 저자들이 하지 않은 비교를 했습니다. 원문 평가 단계는 출판 후 정정 이력과 신고된 이해관계들을 드러냈는데, 4기법 버전에는 이것을 찾아낼 단계 자체가 없었습니다. 그리고 그 시험이 추가 수정을 낳았습니다 — 5단계에 이해관계를 어떻게 제시할지가 들어갔습니다: 사실로 진술하고, 신고된 내용임을 밝히고, 연구의 강점을 같은 호흡에 놓고, 이해관계가 발견에 대한 비판을 대신 지게 하지 말 것. “연구비를 댄 쪽이 이 사안에 입장이 있다”는 계수에 대한 반론이 아니기 때문입니다.

Agent skillScience communicationSource assessmentSelf-verificationClaude Code
14 · Job ApplicationsIt scored a pastry-chef posting at 92%, so now it computes or says it cannot
14
Job Search AI구직 지원 프레임워크
A job application framework that runs on your own machine — and a demo that had to stop pretending before it was worth shipping.자신의 기계에서 도는 구직 지원 프레임워크 — 그리고 내보낼 가치를 갖기 전에 흉내내기를 멈춰야 했던 데모.
The problem

Every application wants the same CV said differently, and the documents involved — your full history, your salary floor, what you will not accept — are among the more private things you own. Hosted résumé tools want them on their servers. But the harder problem turned out to be one the tool created itself: a demo that returns a plausible number is indistinguishable from a tool that works. This one had a dashboard of invented statistics, a badge reading “0% Fabrication”, and a repository containing a fact-grounding rule that forbids inventing numbers. Pasting in a posting for a pastry chef — laminated dough, 5am starts, no computers involved whatsoever — returned a 92% alignment score, a PASS on a salary range nobody had stated, and 100% on Python.지원할 때마다 같은 이력서를 다르게 말해야 하고, 거기 들어가는 문서 — 전체 경력, 최소 희망 연봉, 받아들일 수 없는 조건 — 는 가진 것 중 사적인 편에 속합니다. 호스팅형 이력서 도구들은 그것을 자기 서버에 올리길 원합니다. 그런데 더 어려운 문제는 도구가 스스로 만든 것이었습니다: 그럴듯한 숫자를 돌려주는 데모는 실제로 작동하는 도구와 구별되지 않습니다. 이 앱에는 지어낸 통계로 채운 대시보드와 “0% Fabrication” 배지가 있었고, 같은 레포 안에 숫자를 지어내는 것을 금지하는 fact-grounding 규칙이 있었습니다. 제빵사 구인공고를 붙여넣으면 — 라미네이션 반죽, 새벽 5시 출근, 컴퓨터는 전혀 관여하지 않음 — 92% 적합도와, 아무도 명시하지 않은 급여 범위에 대한 PASS와, Python 100%가 돌아왔습니다.

What it does

The pipeline is a set of Claude Code skills and commands: score a posting, tailor a CV to it, draft a letter, build an interview pack. It reads your documents off your own disk and nothing leaves it, and a fact-grounding protocol governs every output — rephrasing and reordering allowed, inventing a metric or a technology forbidden. The web dashboard is now split honestly. Its evaluator computes: paste a posting and a CV and it reports keyword coverage against a lexicon, a per-category breakdown, a constraint audit that parses salary, work arrangement and years out of the posting text, and — the useful half — the terms the posting asks for that your CV does not contain. All of it deterministic, in the browser, with no API key, which is what lets a static page show a real number. Every other tab is a placeholder and says so on screen, and the notice is rendered once above the tab content so a tab added later cannot ship without it.파이프라인은 Claude Code 스킬과 커맨드 묶음입니다: 공고 평가, 이력서 튜닝, 커버레터 초안, 면접 준비 팩. 사용자 디스크에서 문서를 읽고 아무것도 밖으로 나가지 않으며, fact-grounding 프로토콜이 모든 출력을 통제합니다 — 재표현과 재배열은 허용, 수치나 기술을 지어내는 것은 금지. 웹 대시보드는 이제 정직하게 갈라졌습니다. 평가기는 실제로 계산합니다: 공고와 이력서를 붙여넣으면 어휘집 대비 키워드 커버리지, 카테고리별 분석, 공고 텍스트에서 급여·근무형태·연차를 파싱한 제약 감사, 그리고 — 유용한 절반 — 공고가 요구하는데 이력서에 없는 용어 목록을 냅니다. 전부 결정론적이고, 브라우저 안에서, API 키 없이 — 그것이 정적 페이지가 진짜 숫자를 보여줄 수 있는 이유입니다. 나머지 탭은 placeholder이고 화면에 그렇다고 적혀 있으며, 그 고지는 탭 콘텐츠 위에 한 번만 렌더링되어 나중에 추가되는 탭이 그것 없이 배포될 수 없게 했습니다.

Why it matters

Two rules live in the matching code rather than in the interface, because an interface is where a rule goes to be forgotten. A posting containing no term the lexicon recognises returns no score and says why — the pastry-chef posting is now a test case rather than a 92%. And an unstated constraint is reported as NOT STATED, never as a pass; the third verdict exists because the previous version passed a salary check against a range it had invented. The wording moved with the mechanics: the tool will not say “strong match” or “highly recommended”, and a test asserts it says no such thing, because coverage is a property of two documents rather than a prediction about an interview. What it measures is your document, not you — a skill you have and did not write down scores zero here, exactly as it does in a real screen, which is why the gap list matters more than the score. Twenty-nine tests, gated in CI, no key required to run them — two of them added after the deployed page read “hybrid retrieval” in a fully remote posting as a hybrid work arrangement.두 규칙이 인터페이스가 아니라 매칭 코드 안에 있습니다. 인터페이스는 규칙이 잊히러 가는 곳이기 때문입니다. 어휘집이 인식하는 용어가 하나도 없는 공고는 점수를 내지 않고 이유를 말합니다 — 제빵사 공고는 이제 92%가 아니라 테스트 케이스입니다. 그리고 명시되지 않은 제약은 통과가 아니라 NOT STATED로 보고됩니다. 세 번째 판정값이 존재하는 이유는 이전 버전이 자기가 지어낸 급여 범위에 통과를 줬기 때문입니다. 문구도 함께 움직였습니다: 이 도구는 “strong match”나 “highly recommended”를 말하지 않고, 그런 말을 하지 않는지 검사하는 테스트가 있습니다. 커버리지는 두 문서의 속성이지 면접에 대한 예측이 아니기 때문입니다. 이것이 재는 것은 당신이 아니라 당신의 문서입니다 — 갖고 있지만 적지 않은 스킬은 여기서 0점이고, 실제 서류전형에서도 그렇습니다. 그래서 갭 목록이 점수보다 중요합니다. 테스트 29개, CI에서 게이트, 실행에 키 불필요 — 그중 둘은 배포된 페이지가 완전 재택 공고의 “hybrid retrieval”을 하이브리드 근무형태로 읽은 뒤에 추가됐습니다.

Claude Code skillsIn-browser matchingFact groundingLocal & privateVitest + CI
15 · Clinical ML EvaluationR² of 0.99 on a random split, below zero on new patients
15
Nori vs XGBoost on Clinical Data임상 데이터에서 Nori vs XGBoost
A head-to-head of a training-free tabular foundation model and tuned XGBoost on three healthcare datasets — scored the way a hospital would deploy it, not the way a leaderboard would.학습이 필요 없는 테이블 파운데이션 모델과 튜닝한 XGBoost를 의료 데이터셋 세 개로 맞붙인 비교 — 리더보드 방식이 아니라 병원이 실제로 배포하는 방식으로 채점했습니다.
The problem

Tabular foundation models promise to replace the tuning loop: hand the model some labelled rows and it predicts in one pass, with uncertainty included. My first quick test agreed — Nori beat XGBoost on every split. But that XGBoost was untuned, and the splits were random. Clinical models do not meet random rows in production. They meet new patients, and a benchmark that lets the same patient appear on both sides of the split is measuring whether the model recognises someone, not whether it understands the disease.테이블 파운데이션 모델은 튜닝 과정을 없애 주겠다고 약속합니다: 라벨이 붙은 행 몇 개를 주면 한 번에 예측하고, 불확실성까지 함께 줍니다. 처음 돌린 간단한 테스트도 그렇다고 말했습니다 — Nori가 모든 분할에서 XGBoost를 이겼습니다. 하지만 그 XGBoost는 튜닝되지 않았고, 분할은 무작위였습니다. 임상 모델은 실전에서 무작위 행을 만나지 않습니다. 새로운 환자를 만납니다. 같은 환자가 학습과 테스트 양쪽에 들어가는 벤치마크는 모델이 질병을 이해하는지가 아니라 그 사람을 알아보는지를 재고 있는 것입니다.

What it does

A reproducible notebook compares Nori’s 30-million-parameter checkpoint, untouched and running on a 6 GB laptop GPU, against XGBoost tuned by randomized search on the training fold only, across diabetes progression (442 patients), Parkinson’s severity from voice recordings (5,875 recordings from 42 patients, split by patient), and CMS Medicare payments per discharge. Both models are scored on R², error, the coverage and width of their 90% prediction intervals, and wall-clock cost. A learning curve tests the small-data claim, and a final cell reruns Parkinson’s with a random split to measure exactly how much leakage flatters each model.재현 가능한 노트북 하나가 아무 설정도 건드리지 않은 Nori 3천만 파라미터 체크포인트(6 GB 노트북 GPU에서 실행)와, 학습 폴드 안에서만 랜덤 서치로 튜닝한 XGBoost를 비교합니다. 데이터는 당뇨병 진행도(환자 442명), 음성 녹음으로 추정하는 파킨슨병 중증도(환자 42명의 녹음 5,875건, 환자 단위로 분할), CMS 메디케어 퇴원당 지급액입니다. 두 모델 모두 R², 오차, 90% 예측구간의 커버리지와 폭, 실행 시간으로 채점합니다. 학습곡선으로 적은 데이터에서 강하다는 주장을 검증하고, 마지막 셀은 파킨슨 데이터를 무작위 분할로 다시 돌려 누수가 각 모델을 얼마나 부풀리는지 정확히 잽니다.

Why it matters

The honest result is a split decision. On Medicare payments Nori wins all five folds (R² 0.888 vs 0.872) in about ten seconds a fold against XGBoost’s seventy-six, and on the small diabetes cohort it ties (0.484 vs 0.475) with no tuning at all. Where test rows resemble the context, Nori’s intervals are the better calibrated: 91–92% coverage against XGBoost’s 84–87%, and on Medicare a fifth narrower too. The size of the checkpoint mattered — the six-million-parameter version only tied on Medicare. Then comes the patient split. Neither model predicts a new patient’s severity from voice features — both R² values are negative — but Nori fails harder (−0.71 vs −0.27), and its 90% intervals contain only 22% of the true scores: confidently wrong, exactly where a clinician would trust it. The bigger checkpoint did not fix that. Split by recording instead and the same models score 0.994 and 0.969. Nori gains more from the leak because in-context learning can match a query to near-copies of the same patient. That gap is the finding, not the leaderboard: any evaluation of an in-context model on health data needs a patient-grouped split before anyone believes its score.정직한 결과는 판정이 갈린다는 것입니다. 메디케어 지급액에서는 Nori가 다섯 폴드를 모두 이기고(R² 0.888 vs 0.872) 폴드당 약 10초로 XGBoost의 76초보다 빠르며, 작은 당뇨병 코호트에서는 튜닝 없이 비깁니다(0.484 vs 0.475). 테스트 행이 컨텍스트와 비슷할 때 Nori의 예측구간이 더 잘 보정되어 있습니다: 커버리지 91–92% 대 XGBoost 84–87%, 메디케어에서는 구간 폭도 5분의 1 더 좁습니다. 체크포인트 크기가 중요했습니다 — 6백만 파라미터 버전은 메디케어에서 비기는 데 그쳤습니다. 그다음이 환자 단위 분할입니다. 두 모델 모두 음성 특징으로 새 환자의 중증도를 예측하지 못하고 — R²가 둘 다 음수 — Nori가 더 크게 실패합니다(−0.71 vs −0.27). Nori의 90% 구간은 실제 점수의 22%만 담습니다: 임상의가 믿을 바로 그 지점에서 자신 있게 틀립니다. 더 큰 체크포인트도 이것은 고치지 못했습니다. 녹음 단위로 분할하면 같은 모델들이 0.994와 0.969를 기록합니다. 인컨텍스트 학습은 질의를 같은 환자의 거의 똑같은 행과 짝지을 수 있어서 Nori가 누수의 덕을 더 봅니다. 발견은 리더보드가 아니라 그 격차입니다: 의료 데이터에서 인컨텍스트 모델을 평가한다면, 누구든 점수를 믿기 전에 환자 단위 분할부터 해야 합니다.

Tabular foundation modelsXGBoostPatient-grouped CVPrediction intervalsLeakage auditGPU inference
16 · SecurityCatches the coordinated attack that hides under every per-device threshold
16
Network Traffic Anomaly Detection네트워크 트래픽 이상 탐지
Catches the quiet, coordinated attacks that hide below every per-device alarm.기기별 임계치 밑으로 숨는 ‘조용한 협동 공격’까지 잡아냅니다.
The problem

Security tools watch each machine against a fixed threshold, so a slow, coordinated attack — where every machine stays just under the limit — slips through. And normal daily and weekly traffic swings set off false alarms.보안 도구는 기기마다 고정된 기준으로 감시하기 때문에, 각 기기가 기준을 살짝 밑도는 ‘느리고 협동적인 공격’은 빠져나갑니다. 게다가 평일·주말의 정상적인 트래픽 변동이 오탐을 일으킵니다.

What it does

It learns each entity’s normal rhythm by time-of-week and flags only genuine deviations, so routine ups and downs don’t trigger. Then it adds a “co-spike” check: when many machines in the same network block tick up together, it catches the coordinated attack even though no single one crossed its own line. A classifier labels the type — normal, duplicate, retry, anomalous, DDoS — and every alert ships a structured evidence packet an AI agent can investigate.각 대상의 ‘요일·시간대별 정상 리듬’을 학습해 진짜 이탈만 표시합니다 — 일상적 변동엔 울리지 않습니다. 여기에 ‘동시 급증(co-spike)’ 점검을 더해, 같은 네트워크 대역의 여러 기기가 함께 올라가면 어느 하나도 자기 기준을 넘지 않았어도 협동 공격을 잡아냅니다. 분류기가 유형(정상·중복·재시도·이상·DDoS)을 붙이고, 모든 경보엔 AI 에이전트가 조사할 수 있는 구조화된 증거 묶음이 함께 나갑니다.

Why it matters왜 중요한가

The dangerous attacks are the quiet, coordinated ones. Looking across machines instead of one at a time is what catches them — and the evidence trail keeps a human in the loop.위험한 공격은 조용하고 협동적인 것들입니다. 한 대씩이 아니라 여러 대를 함께 보는 것이 이를 잡는 핵심이고, 증거 기록이 사람의 판단을 끝까지 남깁니다.

Anomaly detectionBehavioral baselinesCross-entity correlationEvidence for agents
17 · Civic DataEvery election number, traceable to its source — and bilingual
17
Electoral Insights Hub선거 인사이트 허브
An open, bilingual analytics platform for South Korean presidential elections and the US midterms — every figure traceable.한국 대선과 미국 중간선거를 위한 공개·이중언어 분석 플랫폼 — 모든 수치가 추적 가능.
The problem

Official election results are scattered across releases, hard to compare from one election to the next, and rarely reproducible — and most dashboards are neither bilingual nor precinct-level, so overseas Koreans, journalists and researchers can’t easily check the numbers themselves.공식 선거 결과는 여기저기 흩어져 있고, 선거끼리 비교하기 어려우며, 재현도 거의 불가능합니다 — 게다가 대부분의 대시보드는 이중언어도, 투표구 단위도 아니어서 재외 한인·기자·연구자가 수치를 직접 확인하기 어렵습니다.

What it does

Two live dashboards. For Korea, an open view of the 18th–21st presidential elections (2012–2025) across all 17 provinces — turnout, vote share, swing, margin and recount — backed by a Gemini + Document AI pipeline that digitizes ~25,000 scanned 21st-election tally sheets (NEC releases and disclosure requests) behind a human-review gate, with a bilingual “Ask the Data” interface that answers with cited numbers and abstains when unsupported. For the US, a live Korean/English midterm forecast with a public calibration report against the Nov 3, 2026 results.두 개의 라이브 대시보드. 한국은 제18–21대 대선(2012–2025)을 17개 시도 전역에 걸쳐 공개 — 투표율·득표율·스윙·격차·재검표 — 하고, Gemini + Document AI 파이프라인이 21대 개표표 약 25,000장(NEC 공개·정보공개)을 사람 검토 게이트 하에 디지털화하며, 이중언어 “Ask the Data”가 인용 수치로 답하고 근거가 없으면 유보합니다. 미국은 한/영 중간선거 예측을 라이브로 제공하고, 2026-11-03 결과 대비 공개 보정 리포트를 냅니다.

Why it matters왜 중요한가

Voters, journalists and researchers get transparent, reproducible election data — every number traceable to its source document, and non-partisan by design.유권자·기자·연구자가 투명하고 재현 가능한 선거 데이터를 얻습니다 — 모든 수치가 출처 문서까지 추적되고, 설계부터 비당파적입니다.

Election analyticsDocument AIBigQueryForecast + calibrationCited bilingual Q&A
18 · AI OperationsFix the cost spike before it hits the invoice
18
Unified Ops AX — self-healing AI fleet telemetryUnified Ops AX — 스스로 고치는 AI 운영 관제
Watches the AI fleet’s cost and errors — and fixes them in milliseconds.AI 운영의 비용·오류를 지켜보다 밀리초 안에 스스로 고칩니다.
The problem

An AI system can run up a huge bill or spiral into errors faster than a human can react. By the time someone notices the cost spike, it’s already on the invoice.AI 시스템은 사람이 반응하기도 전에 큰 비용을 쓰거나 오류로 번질 수 있습니다. 비용 급증을 알아챌 때쯤이면 이미 청구서에 찍혀 있습니다.

What it does

It streams every AI call’s cost, speed, and errors into Splunk in real time. When something spikes — cost over budget, a latency burst, a data-leak rule trip — it automatically switches to a cheaper or safer setting within milliseconds, then restores normal once things calm down. A multi-model router sends each request to the best-value model across providers.모든 AI 호출의 비용·속도·오류를 실시간으로 Splunk에 흘려보냅니다. 비용 초과·지연 급증·데이터 유출 규칙 위반 같은 문제가 생기면 밀리초 안에 더 싸거나 안전한 설정으로 자동 전환하고, 진정되면 원상 복구합니다. 멀티모델 라우터가 각 요청을 제공사 간 ‘가성비 최적’ 모델로 보냅니다.

Why it matters왜 중요한가

At fleet scale a human can’t watch everything. A telemetry-and-remediation loop catches cost runaways before they hit the invoice — and keeps the decisions observable.대규모 운영에선 사람이 전부 지켜볼 수 없습니다. 관측·자동조치 루프가 비용 폭주를 청구서에 닿기 전에 잡고, 그 판단을 관찰 가능하게 남깁니다.

LLM opsSplunk telemetryAuto-remediationMulti-model routing