Before You Convert a Single CA 2E Program, Find Out Which Ones Anyone Still Runs
Sizing a code-generator exit from usage data, one measured pilot and three numbers per program, so the estimate exists before anyone needs it
The first CA 2E program we converted to RPGLE took about five hours to generate and about five more to get compiling. The second started with more than two thousand compile errors, and two weeks later it was down to a few hundred.
Neither number answers the question leadership actually asks: how long to get off the generator entirely? For that you need two more things. First, the list of programs that actually matter. Second, an honest way to scale from a couple of measured pilots to the whole list.
This post is about building both before you need them. The reason to do this early is simple. A plan for leaving a code generator is cheap to build when nothing is wrong. It's very expensive to build in the week something goes wrong with the tool, its licence, its support, or the one person who still knows how to drive it.
The compile mechanics are in article 7, and the terms are in the glossary. This one is about deciding what to convert, and how much it will cost.
The program library is not the list
A long-lived application library has every program anyone ever generated in it. Some replaced versions that were never deleted. Some belong to sites that closed. Some ran once for a data fix in 2014. If you count objects, you size the conversion of the library, not of the application.
The useful list is programs that have run recently, and IBM i already records that. Every object carries usage information: when it was last used, and on how many distinct days. QSYS2.OBJECT_STATISTICS exposes it as a table function, so one query lists every program in a library with its usage:
SELECT objname,
objtext,
objcreated,
change_timestamp,
objsize,
last_used_timestamp,
days_used_count
FROM TABLE (QSYS2.OBJECT_STATISTICS(
OBJECT_SCHEMA => 'PGMLIB',
OBJTYPELIST => '*PGM'))
WHERE objattribute = 'CBL'
AND last_used_timestamp > CURRENT TIMESTAMP - 12 MONTHS
ORDER BY days_used_count DESC;
The objattribute = 'CBL' filter matters if, like us, the generator was still producing COBOL for most of the application. It narrows the list to the programs a model generated as COBOL, the ones a "convert to RPGLE" project is actually about. Article 9 explains why that group is cheaper than hand-written COBOL. Swap the filter for 'RPG' or 'RPGLE' to size the rest.
One small thing that tripped us up: the first version of this query was written against the column's system name, a truncated identifier along the lines of LAST_00001, copied out of a results screen. It worked, but nobody reading it later knew what it meant. Use the long SQL names. They're what the next person will search for.
Usage data has blind spots
The usage columns are good enough to cut a list down. They're not good enough to declare something dead on their own:
The window has to cover the calendar. Our first cut used a window of about seven weeks. That's fine for daily batch and screen programs. It's wrong for anything that runs at month end, quarter end, year end, or once a season. Use at least twelve months, or the list of programs that look dead will include the ones finance needs in January.
A re-created object starts its history over. Usage belongs to the object, not to the name. If a program was recompiled or recreated last month, its usage looks one month old, however long the program has really been in use. Check objcreated next to last_used_timestamp. A recent create date with a short usage history means "unknown," not "rarely used."
Usage counts can be reset. CHGOBJD can reset an object's days-used count, and some shops reset counts as part of housekeeping. Ask before trusting the numbers.
Called-by-name isn't the only way in. A program that's only reached through a menu option nobody uses, or a job scheduler entry that's been on hold for a year, may still be someone's emergency procedure. The usage data tells you who to ask. It doesn't answer for them.
So the output of the query isn't "the list." It's three lists:
- Used recently, often: convert, and these drive the estimate.
- Used, but rarely: confirm the owner and the calendar, then convert or retire.
- Not used in the window: confirm with the business, then archive instead of converting. This group is usually larger than people expect, and it's the cheapest part of any conversion.
Compare against production before you count
Before you trust the list, check that the objects you're counting are the ones running in production. A development or model library can have a newer object, an older object, or a program production never got. Run the same query against the production program library and join on the name. Where the two disagree on create or change timestamps, you've found a program whose source and running object may have drifted. Article 9 covers what to do with those before you regenerate anything.
Measure a pilot, then measure a second one
With the list in hand, you need per-program cost. There's no shortcut: convert a real program end to end and write down what it took.
Track three numbers per program, separately, because they don't move together:
| Number | What it covers | What drives it |
|---|---|---|
| Conversion | Copying the model, changing the target language, generating | Mostly fixed per program once the model copy exists |
| Compile fix | Getting generated source to compile clean | Number of files, access paths and key lists the program touches |
| Testing | Proving the converted program does what the old one did | How easy the inputs are to set up and the outputs to check |
Our first pilot came in at roughly five hours for conversion and five for compile fixes. That's after a false start: the first attempt at the fixes went the wrong way, so the program was regenerated and fixed again from scratch. That's normal. The first program is where you learn the model's conventions and which error messages are causes rather than cascade. Don't extrapolate from it.
The second pilot is the one to calibrate on, and it's worth choosing it deliberately. We picked it for one reason: it was easy to test, and we knew how to get its data. That turned out to be the most useful selection rule we had. A pilot you can't test gives you two of the three numbers and leaves the biggest one blank.
Turning pilots into an estimate
Once two or three pilots are measured, group the recently used programs by the thing that drives effort. File usage is the strongest predictor of compile-fix time, so a rough grouping is enough:
- Light: one or two files, no display file.
- Medium: several files or access paths, batch only.
- Heavy: many files, or any display file. Screen programs cost more to convert, and much more to test.
Assign each pilot to its group, use its measured hours as that group's rate, and multiply. Then write the testing line down as unknown until a pilot has actually been tested. In our own tracking sheet that column still said "TBD" after compile fixes were done. That's the honest state of most conversion estimates, and the one thing you shouldn't hide from whoever reads it.
The result is a table with a row per group: count of programs, hours per program for conversion and compile fixes, and a testing column you fill in as pilots finish. It won't be exact. It will be defensible, because every number in it traces back to a query or a measured program.
What this buys you
On a calm day, this is a few hours of work: one query, two careful pilots and a spreadsheet. On a bad day, when someone asks "what would it take to get off this tool?", it's the difference between an answer by Friday and an answer in a quarter.
It also usually shrinks the project. Archiving programs nobody has run in a year costs almost nothing, and in a library that's been generated into for a long time, that group is rarely small.
The next post covers the decision that shows up as soon as the first compile error is fixed: do you fix it in the model, or in the generated source?
Jaya Krushna Mohapatra is a Warehouse Management Systems Architect focused on enterprise integrations, IBM i modernization, and scalable backend systems.
Top comments (0)